Integrated Inference for Embeddings
July 2024 Product Update

July 2024 Product Update

8/1/2024 · Xian Huang

What this post added

Introduced Pinecone Inference API in public preview, providing low-latency access to embedding and reranking models hosted on Pinecone's infrastructure. This simplifies AI workflows by reducing the need for external tools and infrastructure management. The initial offering includes the multilingual-e5-large model. The post also mentions optimizations using NVIDIA TensorRT and Triton dynamic batching for these inference capabilities.

Read the original post ↗