Managed Knowledge Layer for AI Applications
Benchmarking AI Assistants

Benchmarking AI Assistants

6/25/2024 · Nathan Cordeiro, Roy Miara, Amnon Catav

What this post added

Introduced a new benchmarking scheme for AI assistants, including a novel 'Answer Alignment Score' metric based on correctness and completeness. This metric aims to better correlate with human judgment than existing unsupervised metrics by evaluating factual entailment, contradiction, and neutrality against ground truth. The post details the protocol for calculating this score and presents results comparing Pinecone Assistant against OpenAI assistants on three benchmark datasets (FinanceBench, Open Australian Legal, NQ-HARD), demonstrating Pinecone Assistant's superior performance.

Read the original post ↗