For Lumen, we built an evaluation harness that runs real business questions through the assistant, executes the generated query against a real data source, and checks that the answer is actually correct - then scores every leading model on accuracy, cost, and speed, side by side.
**Learning**
Accuracy came from context, not the priciest model.
In our AI data assistant, every answer is measured against ground truth, ~1 cent per question, and zero data retention.
Try now.

