HomeMicrosoft finds costs multiply during some AI model upgradesUncategorizedMicrosoft finds costs multiply during some AI model upgrades

Microsoft finds costs multiply during some AI model upgrades

Microsoft has found that developers upgrading to some new AI models face unpredictable token consumption and escalating costs.

The company recently evaluated AI agent execution across two of Anthropic’s major model versions, Claude Sonnet 4.6 and Claude Sonnet 5, deploying GitHub Copilot Chat within Visual Studio Code on Windows.

The assessment tested 150 specific agent tasks divided across 15 technical scenarios. To ensure objective measurement, engineers executed five runs per model per scenario. Each individual run faced evaluation against binary criteria organised into a Select gate and specific quality dimensions, scored by an LLM judge calibrated for consistency. Costs were computed directly from actual per-turn token data priced against GitHub Copilot rates.

These scenarios isolated two distinct categories of developer workloads. Engineers tested Azure architecture design grounded heavily in Microsoft Learn documentation alongside complex SharePoint Framework project upgrades.

Rate cards favour the newer model by displaying a 33 percent price reduction across every token category. Input tokens for Sonnet 5 cost $2 per million compared to $3 for the previous iteration. Cached input dropped to $0.20 from $0.30, and output token pricing fell from $15 to $10 per million.

Hidden token spikes

Billing structures rely on active token consumption rather than static rate cards. Sonnet 5 consumes a vastly higher volume of tokens to execute the exact same engineering directives.

The financial impact diverges based on the technical workload being processed. Microsoft engineers ran 12 scenarios, conducting 60 distinct runs per model. Sonnet 5 consumed 12 times more tokens at the median during these architecture tasks. One scenario recorded a single run burning 47 times the typical baseline token volume.

Cost reductions occasionally surface despite the increased volume. The newer model averaged $0.47 per run on these tasks against $0.54 for Sonnet 4.6. The token consumption increase remained moderate enough in this specific context for the 33 percent rate discount to generate a 12 percent cost reduction.

Budgeting requires predictability, which the newer model architecture fails to provide. Sonnet 4.6 maintained tight operational clustering during straightforward architecture tasks, with most runs consuming between 14,000 and 45,000 tokens. Sonnet 5 exhibited extreme and erratic variance.

Degradation in idiomatic code generation

Cost reductions hold little professional value without maintaining baseline output quality. Both models completed the tasks at a 75 percent success rate on the ‘Select’ gate, meaning the agent successfully attempted the requested work.

Output quality demonstrated a regression in the newer model. Sonnet 4.6 achieved a 90 percent score on the ‘Idiomatic’ dimension across the nine scenarios where both models produced usable output. This dimension evaluates whether the generated output follows established industry patterns and conventions. Sonnet 5 scored 78 percent on the same metric. The older iteration matched or outperformed the newer version in eight of those nine comparative scenarios.

Engineers designing an IoT analytics architecture witnessed this performance gap. Both models completed the assignment in every attempt, but Sonnet 4.6 passed idiomatic checks in four out of five runs. Sonnet 5 passed the exact same check only once using the identical prompt. The newer model produced measurably worse output while consuming more tokens across the majority of scenarios.

The IoT analytics scenario saw one Sonnet 5 run consume 16,000 tokens while another executed on the same baseline prompt burned 6.6 million tokens. Outlier consumption events occurred occasionally on the older model but became the standard operating pattern for the newer iteration.

Code migration and instruction adherence

While architectural design revealed a degradation in output quality for the newer model, shifting the workload to complex codebase upgrades completely reversed the performance and reliability dynamics.

The Microsoft evaluation tested three specific SharePoint Framework project upgrade scenarios. These tasks included migrating a build system from gulp to Heft and updating a legacy ESLint configuration to a flat config. Sonnet 5 passed the Select gate in 100 percent of the runs, dominating the 60 percent success rate achieved by Sonnet 4.6.

Upgrading from SPFx version 1.21.1 to 1.22.0 highlighted the instruction adherence capabilities of the newer release. Sonnet 4.6 failed all five execution attempts by actively overriding the direct user instruction and adopting version 1.22.1 based entirely on its Microsoft Learn grounding context. Sonnet 5 executed the version instruction precisely during every single attempt.

This enhanced instruction adherence introduced financial penalties at the execution layer. The gap in token consumption hit a factor of 10 during the 15 code upgrade runs per model. Sonnet 5 cost $2.01 per run, making it 3.7 times more expensive than the $0.55 median cost of Sonnet 4.6.

Deep execution attempts drove this severe financial variance, highlighted by an extreme outlier that underscores the unpredictability of newer agents. One Sonnet 5 run consumed a staggering 69 million tokens by conducting extensive web fetching to discover undocumented migration steps. This expansive run met 21 out of 30 strict evaluation criteria. The behaviour remains strictly irreproducible in enterprise environments, however, as four out of five runs in each scenario failed to achieve this analytical depth.

Undocumented environments limit agent capabilities

Configuration correctness remained frozen at zero percent across all SharePoint Framework scenarios for both assessed models. Neither model successfully navigated the structural toolchain migrations involving gulp, Heft, or ESLint.

Structural migrations require executing specialised steps rarely centralised in official documentation. These actions include altering build tool flags, restructuring package manifests, deleting deprecated files, and migrating entire configuration formats.

The Microsoft engineering team identified seven missing configuration changes required for task completion. Generative agents cannot apply undocumented processes, rendering the selected model version entirely irrelevant when the underlying grounding content contains omissions.

See also: Google Cloud details full-stack AI architecture for developers

AI & Big Data Expo banner by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

Developer is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

Home
Services
Careers
Call Us
Contact