Alibaba.com Tests 13 AI Models: Best Completes Only 61.7% of Real Commerce Tasks
Alibaba.com's CommerceAgentBench reveals that even the top-performing AI model, Claude Opus 5, fails to complete nearly 40% of real-world trade tasks, with the biggest gaps in landed-cost calculations, return disputes, and multi-leg shipping routes.

Alibaba.com Tests 13 AI Models: Best Completes Only 61.7% of Real Commerce Tasks
Alibaba.com has released a public testing tool called CommerceAgentBench on GitHub, scoring 13 AI model families on 107 real-world commerce tasks, with results announced in September 2026. The test uses a pass/fail system, awarding points only when a task is completed correctly—for example, a product must be listed with the correct attributes, and a shipment must be booked on a route that actually exists.
The highest scorer was Claude Opus 5, with a task completion rate of 61.7%. Alibaba.com President Zhang Kuo wrote in Fortune on September 9 that this score is "high enough to be useful, and low enough to be a warning." The test covered procurement, logistics, product listing, fulfillment, and after-sales service.
According to the test results, AI agents performed worst on tasks requiring information retention across multiple steps, such as landed-cost calculations, return disputes, and multi-leg shipping routes. Separately, according to PYMNTS Intelligence data, approximately 132 million U.S. adults have purchased retail goods with AI assistance, with Amazon accounting for 59% of such purchases.
Source: https://www.pymnts.com/news/artificial-intelligence/2026/commerce-test-shows-where-ai-agents-break-down/
Provenance & status
- Byline
- OceanAlt Editorial
- First published
- 2026-09-19
- Last updated
- 2026-09-19
- Content type
- Newsflash
- Source material
- View original ↗
Related reading

Visa and Mastercard Join Ant International on a KYA Interoperability Framework as Agent Identity Standards Begin to Converge

Félix Raises $200M Led by a16z: Stablecoin Infrastructure Shifts from Remittances to Agent Economy Settlement
U.S. Congress Holds First Hearing on AI Agent Payment Rules: Authorization, Settlement, and Identity Take Center Stage
Paste a payee address before you pay and see whether it's on a sanctions list, through a mixer, or tagged for fraud.
This judgement can sit inside your own product
One line of code; it touches neither your CSS nor your JS. The same pre-settlement judgement can appear in your articles, on your wallet's confirmation screen, or as an endpoint your agent calls before it pays.

