Benchmarking Structured Conversation Extraction Across Three Evaluation Layers: A Comparison of Eight Models with a Public Replication
Evaluating structured extraction from conversational data with a single quality metric conceals the failure modes that matter most to systems built on top of the extraction. This paper presents a benchmark that evaluates extraction quality across three layers: field comparison against a reference extraction, evaluation by a panel of three judge models reading the raw conversation, and propagation ...