Kafka Broker and Streams Enhancements
Why Scrapinghub’s AutoExtract Chose Confluent Cloud

Why Scrapinghub’s AutoExtract Chose Confluent Cloud

10/3/2019 · Ian Duffy

What this post added

Scrapinghub's AutoExtract service uses Confluent Cloud for its Kafka needs to scale its AI data extraction API. They partition URLs for fetching, rendering, and screenshotting using Kafka, and distribute resource-intensive AI data extraction tasks across multiple instances. The move to Confluent Cloud allowed them to offload Kafka infrastructure management, providing scalability and cost benefits. They detail their evaluation of alternatives like self-hosting Kafka on Kubernetes and Amazon MSK, ultimately choosing Confluent Cloud for its managed nature, consumption-based pricing, and vendor independence. They also describe initial setup, load testing, and workarounds for missing tools like Burrow by using a consumer metrics exporter.

Read the original post ↗