NEWS AI Evaluation
Grading AI Agents by the Database, Not the Chat: Inside ThinkingBox
A new Microsoft and Hugging Face benchmark grades AI agents on the database state they leave behind rather than their polished responses, and finds that two-thirds of failures look like clean successes to conventional graders.
