Ashita Orbis

Research

Working papers, analyses, and technical investigations. Not peer-reviewed.

Benchmarking Structured Conversation Extraction Across Three Evaluation Layers: A Comparison of Eight Models with a Public Replication

Ashita Orbis | June 8, 2026 | 21 min read | working-paper

Evaluating structured extraction from conversational data with a single quality metric conceals the failure modes that matter most to systems built on top of the extraction. This paper presents a benchmark that evaluates extraction quality across three layers: field comparison against a reference extraction, evaluation by a panel of three judge models reading the raw conversation, and propagation ...

The Etymology Tax: Etymological Register Effects on LLM Multi-Step Reasoning

Ashita Orbis | March 6, 2026 | 42 min read | working-paper

This study tests whether the etymological register of input text affects large language model performance on multi-step reasoning tasks. Using 250 murder mystery narratives from the MuSR benchmark, eight models were evaluated across six conditions in a partially crossed design varying vocabulary register (baseline, Germanic, Latinate) and clarification, with explicit clarification applied only to ...

Convergent AI-Mediated Personality Assessment: Psychometric Profiling and Narrative Inference from Digital Communication Data

Ashita Orbis | March 3, 2026 | 96 min read | working-paper

This paper examines the convergence between two AI-mediated approaches to personality assessment applied to a single participant: explicit psychometric measurement through seventeen validated instruments triangulated across three inference methods, and implicit personality inference through literary narrative generation from a 267MB text messaging archive. The explicit approach (Psyche) produced a...