New Open Source Benchmark Scores AI Agents on Their Ability to Learn and Perform Complex Actions

ⓘ This article is third-party content and does not represent the views of this site. We make no guarantees regarding its accuracy or completeness.

SAN FRANCISCO, Sept. 16, 2026 (GLOBE NEWSWIRE) -- There is now a way to measure how well an AI agent is able to learn and take action on the job it was built to do. The Agent Effectiveness Index (AEI), released today as a free and open-source benchmark, scores and ranks AI agents on their ability to understand complex, real-world processes, take proactive actions, and keep learning without drifting as processes change. It was built by Brackett, which has also launched its Connected Agentic Workforce platform.

Organizations are rushing to build AI agents, but they have no standard way to evaluate whether those agents will perform in production. As agents are increasingly tasked with inference and applying context, this includes whether they are able to make judgment calls, exceptions, and escalation patterns that make a process reliable, or if they're producing inconsistent results at a high cost. It's the equivalent of hiring a company's worth of people with no way to evaluate performance. The Agent Effectiveness Index addresses that gap.

“Most evaluations for AI agents focus on static knowledge, or what it knows. But this doesn't tell you whether or not that agent can complete the task it was created for, because real world tasks require things like judgement calls, exceptions, or knowing when to bring in a human. These are learned from experiences, not training data. Until now there has been no way to score that capability,” said Ehsan Azarnasab, co-founder and Chief Scientist of Brackett and formerly Principal Scientist on Microsoft’s GenAI Platform team.

AEI evaluates agent systems across three dimensions:

  • Business Understanding: whether an agent grasps how a specific company actually works, and grounds its answers in real evidence rather than plausible guesses.
  • Operational Execution: whether an agent produces correct results, handles exceptions, and stays inside the authority it was given.
  • Learning Persistence: whether teaching an agent genuinely changed its behavior, and whether that holds on new cases, after time passes, and when the rules change.

The Index publishes its first scores today, measuring learning and comprehension across three agent systems evaluated on the same demonstration: Brackett, OpenAI's Codex, and Anthropic's Claude. The full task set, scoring code, and methodology are available on Brackett's GitHub under the MIT License, with scoring for execution, transfer, and retention to follow as the Index expands toward a complete picture of agent effectiveness.

"The next era of work isn't humans versus agents, it's humans and agents becoming genuinely better together," said Jaideep Sarkar, Co-Founder and CEO of Brackett. "We built Brackett because we’ve seen a recurring gap in which companies try to automate tasks, but without a way to build compounding, connected intelligence that they actually own. It's also why we're opening the Agent Effectiveness Index to the world. You can't build a trustworthy agentic workforce without a way to measure it."

Also live today is Brackett’s Connected Agentic Workforce Platform. Using its Capture, Codify, and Compound methodology, Brackett turns simple conversations, with no code required, into agents that learn how to execute complex processes. Agents then progress to running tasks at scale with real consistency and control, and then to acting on the judgment they've learned over time. Brackett connects these agents to each other, to the systems they run in, and to the people who trained them, so every workflow makes the next one smarter across the business. The intelligence the organization builds along the way is retained as something it owns rather than something it rents from a model vendor.

The Agent Effectiveness Index is available now on Github, free and open source. To learn more about Brackett's Connect Agentic Workforce Platform for Enterprises, book a demo here.

About Brackett

Brackett is the Connected Agentic Workforce platform. It enables enterprises to capture how their people actually work, codify it into agents that learn real organizational judgment, and compound that intelligence into an asset the company owns. Founded by former technical leaders from Microsoft, Rubrik, and Amazon, Brackett is headquartered in San Francisco and is backed by Focal and Heavybit. Learn more at www.brackett.ai.


Media Contact
Jennifer Lankford
Lankford Communications
jennifer@lankfordpr.com

Primary Logo

Report this content

If you believe this article contains misleading, harmful, or spam content, please let us know.

Report this article

More News

View More

Recent Quotes

View More
Symbol Price Change (%)
AMZN  247.64
-0.78 (-0.31%)
AAPL  333.29
+1.95 (0.59%)
AMD  522.01
+17.81 (3.53%)
BAC  58.59
-0.94 (-1.57%)
GOOG  341.41
-0.02 (-0.01%)
META  674.38
+4.14 (0.62%)
MSFT  493.45
-3.67 (-0.74%)
NVDA  214.88
+2.71 (1.28%)
ORCL  142.80
+2.45 (1.75%)
TSLA  360.48
+3.90 (1.09%)
Stock Quote API & Stock News API supplied by www.cloudquote.io
Quotes delayed at least 20 minutes.
By accessing this page, you agree to the Privacy Policy and Terms Of Service.