6/25/2024 · Nathan Cordeiro, Roy Miara, Amnon Catav
What this post added
Introduced a new benchmarking scheme for AI assistants, including a novel 'Answer Alignment Score' metric based on correctness and completeness. This metric aims to better correlate with human judgment than existing unsupervised metrics by evaluating factual entailment, contradiction, and neutrality against ground truth. The post details the protocol for calculating this score and presents results comparing Pinecone Assistant against OpenAI assistants on three benchmark datasets (FinanceBench, Open Australian Legal, NQ-HARD), demonstrating Pinecone Assistant's superior performance.