
9/12/2025 · Daniel Azoulai
What this post added
Tavus implemented Qdrant Edge for their conversational AI, colocating per-conversation edge vector stores with conversational workers. This removed network latency, reducing retrieval time to 20-25ms and enabling end-to-end utterance-to-utterance timing near 500-600ms. The architecture allows for retrieval on every utterance and supports grounding every turn with private knowledge, improving conversational quality and simplifying onboarding for customers.