
10/3/2019 · Ian Duffy
What this post added
Scrapinghub's AutoExtract service uses Confluent Cloud for its Kafka needs to scale its AI data extraction API. They partition URLs for fetching, rendering, and screenshotting using Kafka, and distribute resource-intensive AI data extraction tasks across multiple instances. The move to Confluent Cloud allowed them to offload Kafka infrastructure management, providing scalability and cost benefits. They detail their evaluation of alternatives like self-hosting Kafka on Kubernetes and Amazon MSK, ultimately choosing Confluent Cloud for its managed nature, consumption-based pricing, and vendor independence. They also describe initial setup, load testing, and workarounds for missing tools like Burrow by using a consumer metrics exporter.