8/1/2024 · Xian Huang
What this post added
Introduced Pinecone Inference API in public preview, providing low-latency access to embedding and reranking models hosted on Pinecone's infrastructure. This simplifies AI workflows by reducing the need for external tools and infrastructure management. The initial offering includes the multilingual-e5-large model. The post also mentions optimizations using NVIDIA TensorRT and Triton dynamic batching for these inference capabilities.