
6/3/2026
What this post added
This post details the application of the hybrid worker-advisor AI architecture to the Legal Agent Benchmark. It demonstrates how a GLM 5.1 worker can selectively invoke Claude Opus 4.7 as an advisor for specific sub-tasks, achieving a higher all-pass rate (18/100) at a lower cost ($368) compared to using Claude Opus 4.7 end-to-end ($954). It also details the results of supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT) of Kimi K2.6 on Legal Agent Benchmark trajectories, showing improvements in all-pass rate and mean score. The post emphasizes the integration of these techniques on the Fireworks platform for seamless experimentation and production deployment.