OceanAltOceanAlt
Agent Economy2026-09-19Event 2026-09-182 min read

Alibaba.com Tests 13 AI Models: Best Completes Only 61.7% of Real Commerce Tasks

Alibaba.com's CommerceAgentBench reveals that even the top-performing AI model, Claude Opus 5, fails to complete nearly 40% of real-world trade tasks, with the biggest gaps in landed-cost calculations, return disputes, and multi-leg shipping routes.

OOceanAlt EditorialSource

Alibaba.com Tests 13 AI Models: Best Completes Only 61.7% of Real Commerce Tasks

Alibaba.com has released a public testing tool called CommerceAgentBench on GitHub, scoring 13 AI model families on 107 real-world commerce tasks, with results announced in September 2026. The test uses a pass/fail system, awarding points only when a task is completed correctly—for example, a product must be listed with the correct attributes, and a shipment must be booked on a route that actually exists.

The highest scorer was Claude Opus 5, with a task completion rate of 61.7%. Alibaba.com President Zhang Kuo wrote in Fortune on September 9 that this score is "high enough to be useful, and low enough to be a warning." The test covered procurement, logistics, product listing, fulfillment, and after-sales service.

According to the test results, AI agents performed worst on tasks requiring information retention across multiple steps, such as landed-cost calculations, return disputes, and multi-leg shipping routes. Separately, according to PYMNTS Intelligence data, approximately 132 million U.S. adults have purchased retail goods with AI assistance, with Amazon accounting for 59% of such purchases.

Source: https://www.pymnts.com/news/artificial-intelligence/2026/commerce-test-shows-where-ai-agents-break-down/

Provenance & status

Byline
OceanAlt Editorial
First published
2026-09-19
Last updated
2026-09-19
Content type
Newsflash
Source material
View original ↗

Cite this piece

OceanAlt Editorial (2026). "Alibaba.com Tests 13 AI Models: Best Completes Only 61.7% of Real Commerce Tasks". OceanAlt. https://oceanalt.com/en/articles/flash-auto-mu7dumqi-hlnu (accessed 2026-09-19)

This piece follows our editorial and fact-checking standards. Found an error? tell us. Once verified, the correction will be published right here.

TRY IT · FREE, NO SIGNUP

Paste a payee address before you pay and see whether it's on a sanctions list, through a mixer, or tagged for fraud.

This judgement can sit inside your own product

One line of code; it touches neither your CSS nor your JS. The same pre-settlement judgement can appear in your articles, on your wallet's confirmation screen, or as an endpoint your agent calls before it pays.

The widget collects no reader identity. Integrating does not mean OceanAlt endorses your product, or any address on your page.