
9/27/2023 · Rita Kozlov, James Allworth, Seph Zdarko
What this post added
This post announces the launch of Workers AI for serverless GPU inference on Cloudflare's global network, Vectorize as a vector database for storing embeddings, and AI Gateway for caching, rate limiting, and observing AI deployments. It also highlights partnerships with NVIDIA, Microsoft, Hugging Face, Databricks, and Meta, and discusses the strategic placement of inference workloads on the network edge. The post details the use of ONNX runtime for seamless model execution across devices, edge, and cloud, and the integration of Llama 2 and Hugging Face models.