The AI model matters less than the harness wrapped around it, according to a growing chorus of CIOs

Started by Jonathan, Today at 12:31 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: The AI model matters less than the harness wrapped around it, according to a growing chorus of CIOs   Views(Read 63 times)
Active members in this topic:
Jonathan(1)

Jonathan

A recent piece in the Wall Street Journal's CIO Journal lays out a shift happening quietly inside enterprise AI teams, the model powering a given AI product increasingly matters less than the software layer wrapped around it, commonly called the harness. That harness includes things like system prompts, tool definitions, memory, permission policies, and the orchestration logic that actually lets an AI agent do useful work rather than just answer a single question.

Supporting research from AI consultancy Systima found something that should worry any CIO comparing vendors purely by model benchmark scores. Running the exact same underlying model, Claude Sonnet 4.5 in this case, through two different harnesses, Anthropic's own Claude Code versus the open source OpenCode, produced sharply different token overhead purely because of differences in how each harness was configured. Same model, wildly different cost and behavior, entirely because of the scaffolding around it.

This reframes a lot of what enterprises should actually be evaluating when they pick an AI vendor. Comparing raw model capability between providers increasingly misses the point if the harness quality varies just as much or more between competing products built on the exact same underlying model. A well designed harness can make a mediocre model perform like a great one on a specific task, and a poorly designed harness can waste enormous amounts of money and effort even wrapped around the best model available.

The competitive implications go further than efficiency too. Companies with harnesses deeply wired into a customer's own data, permissions, and workflows create a kind of lock in that has nothing to do with which underlying model is technically the smartest. That's arguably why Anthropic's enterprise revenue keeps climbing even in periods where its raw benchmark scores aren't the industry's best, the thing customers actually depend on day to day is the surrounding system, not the isolated model score sitting at the center of a leaderboard

GG no re

Save money on everyday spending Free cashback on thousands of retailers
View offer