Brackett opens AI agent benchmark to close workforce evaluation gap

A new open-source index scores AI agents on real-world judgment, threatening to expose the performance gap in enterprise agentic deployments.

An aerial view shows a campus of numerous identical modern white and glass buildings, some connected by elevated glass bridges, set amidst manicured green lawns and concrete pathways under bright daylight.

Brackett, a San Francisco-based startup founded by former technical leaders from Microsoft, Rubrik and Amazon, has released the Agent Effectiveness Index (AEI), a free, open-source benchmark designed to score AI agents on their ability to learn organisational processes, handle exceptions, and retain that learning as conditions change. Launched alongside Brackett's own Connected Agentic Workforce platform, the index publishes its first comparative scores today, pitting Brackett's own system against OpenAI's Codex and Anthropic's Claude on a standardised demonstration task.

The timing is deliberate. Enterprise adoption of agentic AI has accelerated sharply through 2025 and into 2026, with McKinsey and Gartner both tracking a surge in production deployments. Yet the field has lacked a shared performance yardstick that goes beyond static knowledge tests. Most existing evaluations, as Brackett's co-founder and chief scientist Ehsan Azarnasab notes, measure what a model knows rather than whether it can execute the task it was built for. "Real world tasks require things like judgement calls, exceptions, or knowing when to bring in a human," he said. "These are learned from experiences, not training data."

Three dimensions, one index

AEI evaluates agent systems across three axes. Business Understanding tests whether an agent grounds its answers in a specific company's actual operating reality, rather than producing plausible but generic responses. Operational Execution assesses whether it produces correct outputs, manages edge cases, and stays within the authority it has been granted. Learning Persistence measures whether the agent genuinely updates its behaviour after being taught, and whether those updates hold across new cases, time gaps, and rule changes. The full task set, scoring code and methodology are published on GitHub under the MIT Licence, with execution, transfer and retention scoring to follow.

The platform Brackett launched in parallel takes a no-code approach: operators describe their workflows conversationally, and the system codifies those descriptions into agents that the organisation retains as an owned asset rather than a licensed capability. CEO Jaideep Sarkar frames the strategic logic around compounding intelligence: "Companies try to automate tasks, but without a way to build compounding, connected intelligence that they actually own."

Wider implications for the agentic workforce market

The AEI's open-source release is a calculated market-positioning move, but it carries genuine cross-sector significance. The absence of standardised agent evaluation has been a meaningful brake on enterprise AI adoption in regulated industries: financial services firms, healthcare operators and defence contractors all face governance obligations that make it difficult to deploy systems whose performance cannot be independently audited. A credible, vendor-neutral benchmark lowers that barrier, and could accelerate deployment in sectors where the productivity case is strongest but the compliance burden has been highest.

There is also a capital-allocation dimension worth noting. The agentic AI infrastructure layer is attracting substantial venture interest, with firms including Andreessen Horowitz, Sequoia and Coatue allocating to workflow-automation and AI-agent platforms through 2025. Brackett itself is backed by Focal and Heavybit. An open benchmark that separates genuine operational performance from marketing claims stands to shift capital towards agents that demonstrably learn and persist, and away from those that score well on static capability tests but degrade in production. That dynamic matters to the growing cohort of corporate venture arms and institutional investors running AI-adoption programmes across portfolio companies simultaneously.

The index also puts competitive pressure on the hyperscalers whose agent offerings appear in the first published scoring table. For OpenAI and Anthropic, appearing in a third-party benchmark they did not commission is a reputational double-edged proposition: good scores validate enterprise positioning; poor ones, or even middling ones, become a reference point that prospects and procurement teams will cite. Whether the methodology withstands scrutiny from the broader research community, given that it was built by a commercial participant in the same market, will be the index's first real test of credibility.