
7/17/2025
What this post added
This post introduces FutureBench, a new benchmark for evaluating AI agents' ability to predict future events. It details the methodology for generating prediction questions from news and prediction markets, the technical stack used (DeepSeek-V3, Firecrawl, Tavily), and a three-level evaluation framework (framework comparison, tool performance, model capabilities). The post also presents initial results comparing different LLMs (GPT-4.1, Claude3.7, DeepSeek-V3) and analyzes their action patterns and prediction strategies.