Consumer AI has developed a new trick. It can offer expert help and deliver the work of an intern without changing its tone of voice.
The answer still arrives in polished English. It is confident, organised and reassuringly quick. What changes from one question to the next is the amount of intelligence, context and factchecking behind it.
This is easily dismissed as the usual complaint that a technology was better before its manufacturer started improving it. There is more to it this time. The major AI companies openly manage how much computing effort individual questions receive. Prompts can be routed between models, reasoning budgets can change, older context can be compressed and smaller models can take over when usage allowances run out.
In short, the AI you were using six months ago might not be the same AI you are using today, even if it carries the same logo and you changed nothing.
What enshittification actually means
Cory Doctorow coined “enshittification” to describe the way digital platforms decay as they change who receives the value.
The platform begins by treating users exceptionally well because it needs them. Once enough people depend on it, the balance shifts towards the businesses paying to reach those users. When both sides are locked in, the platform takes more for itself. Users receive a worse service, business customers pay more for less access, and leaving is difficult because everybody else is still there.
It is more specific than a product simply becoming worse. Enshittification depends on control. The platform sits between two groups, changes the rules and keeps increasing its share of whatever passes between them.
Google Search is a clean example
Early Google replaced cluttered web directories and mediocre search engines with a blank page and remarkably good organic results. Type the question, find the relevant website and leave. That simplicity attracted users, and the users attracted advertisers. Publishers then reorganised much of the web around Google because appearing in its results could build an audience or a business.
The results page gradually acquired more advertising, shopping results, maps, featured snippets, knowledge panels, videos and other Google-controlled features. Organic links remained, but paid placements and Google-controlled material often occupied more of the page above them.
Google’s AI Overviews move the same process another step forward. Google can use information gathered from external websites to compose an answer before the user visits any of them. Google says its AI results include prominent links and help people explore the web.
Publishers see the other side of that transaction. A 2026 randomised field experiment involving 1,065 US desktop Chrome users found that AI Overviews reduced outbound organic clicks by about 38% on queries where they appeared. Searches ending without a click rose from 54% to 72%, while sponsored clicks did not materially change.
The searcher may receive a faster answer. Google retains the searcher and the advertising opportunity. However, the publisher that supplied the information has a smaller chance of receiving the visit that helps pay for creating it.
Is enshittification already creeping into consumer AI?
AI providers started by giving consumers access to systems that felt implausibly capable for the price; sometimes for no price at all. Those systems are now embedded in research, writing, coding, customer service, marketing and business planning. People are building workflows around them.
The expensive capability has not disappeared. Access to it is being divided into tiers of speed, effort and reliability.
OpenAI has described ChatGPT as a routed system that decides whether a prompt should receive a fast response or deeper reasoning. It has also said that mini models can handle requests after specified usage limits are reached. Google gives Gemini developers direct control over thinking levels and recommends lower settings where latency and token consumption matter. Anthropic offers effort controls for Claude and says lower effort responds faster while consuming rate limits more slowly.
These are sensible engineering decisions. They are cost controls as well.
Different grades of service are normal. Print companies can sell draft, production and high-quality output at different prices. The problem starts when every job is sold under one specification and the pressroom quietly decides how much quality each file deserves, often without considering either context or customer expectation.
That is increasingly how consumer AI feels. The user sees one service but may receive a lower grade allocation of it.
A basic level AI service is perfectly capable of extracting an address, reformatting a table or shortening three paragraphs. Give it a decision involving competing evidence, missing information and commercial risk, and its instinct is more likely to be rapid completion than thorough investigation. It knows what a good answer looks like, but it may not have done the work required to reach one.
The intern is wearing the expert’s suit
The dangerous AI answer is rarely gibberish. Gibberish is easy to reject. The dangerous answer is coherent, plausible and slightly wrong in ways that often require more expertise to detect than the answer required to produce.
A simplified model can write with much the same confidence as an expensive reasoning model. It knows where the headings go, how a qualification should sound and how to finish with a sensible recommendation. The presentation survives even when the reasoning underneath it has been cut short.
The weakness appears in the work. The model answers the general subject instead of the precise question. It fills an evidential gap with the most likely explanation. It forgets a constraint established earlier in the conversation. It claims to have checked something when it has inferred it. Asked to change two elements in an image or document, it changes six because it decided to reconstruct the task rather than follow it.
What I am increasingly seeing is subtler. Ask it to research something and it suggests additional research that I should do myself. Unless challenged, the task quietly changes from doing the work to advising me how I might do it.
And when it is challenged, it always manages to explain these mistakes beautifully.
Anyone using AI professionally has probably seen the increasing inconsistency. The same product can be startlingly capable on Monday and oddly literal, forgetful or superficial on Tuesday. Generative models are inherently variable, so routing cannot be blamed for every poor answer.
Model updates, system instructions, context compression and tool failures all contribute. Providers reveal too little for an outsider to identify the cause of any individual failure with confidence.
From the user’s side, the diagnosis matters less than the result. The service cannot be assumed to deliver the same grade of assistance every time.

What this might look like in a print business
Consider a print company asking exactly the same question of two different AI service tiers:
Q. We produce 18,000 square metres a month using two 3.2 metre printers. Each averages 55 saleable square metres per hour on our normal work, and the department runs for 160 scheduled hours each month. A manufacturer is offering us one replacement printer advertised at 300 square metres per hour for £240,000. Should we replace both existing machines?
A properly provisioned reasoning model should notice that the figures cannot yet support the decision. Its answer might read:
A. Your existing printers produce a combined 110 saleable square metres per hour. The proposed machine’s 300 square metre figure is an advertised maximum, so it cannot be compared directly with your measured production figure.
Ask the supplier to demonstrate the printer using your normal media, resolution, ink coverage and quality requirements. You also need its expected uptime, ink consumption, service costs and operator requirements, together with the capacity of your finishing department.
Across 160 scheduled hours, the replacement would need to sustain 112.5 saleable square metres per hour before allowing for any downtime. At 85% uptime, it would need approximately 132 square metres per hour merely to produce your current monthly volume. Providing 20% capacity above that demand would require approximately 159 square metres per hour at the same uptime.
Replacing two printers with one also removes redundancy. A breakdown would stop all roll-to-roll production. The £240,000 investment cannot be justified from the advertised speed alone.
The valuable part of that answer is its refusal to decide. It separates a measured production rate from a headline specification, calculates the minimum usable output and identifies the single point of failure created by replacing two printers with one.
What if the service routes the question to a faster, cheaper AI setting? The answer may still look convincing:
A. Yes. Your existing printers produce a combined 110 square metres per hour, while the new printer produces 300 square metres per hour. This represents a 173% increase in productivity. Your monthly workload could be completed in 60 hours instead of approximately 164 hours, reducing labour, energy and maintenance costs. Replacing the two older machines therefore appears to be a sound investment.
The maths is correct. The answer is wrong.
It compares 110 square metres of measured, saleable production with a 300 square metre headline speed that may apply only to a faster, lower-quality mode. Labour, energy and maintenance savings are invented because the prompt supplied none of those costs. The answer ignores uptime, finishing capacity and the loss of production redundancy.
Most seriously, it turns missing evidence into a £240,000 recommendation.
The user thinks the question has received expert analysis because the response is fluent, specific and contains calculations. What arrived was intern-level reasoning delivered with senior-level confidence. The error may not become visible until the new printer is installed, and the old machines have gone.
Worse still, the same service may deliver something close to the first answer today and the second answer a few days later. A model update, routing decision or policy change may be responsible, but nothing visible to the user necessarily explains the difference.
Which services are most exposed?
Free tiers, mini models and products carrying labels such as Flash, Flash-Lite or Haiku are designed around speed and economy. That is useful product segmentation when the choice is visible, and the task suits the model.
Default and Auto modes are harder to judge because capability can vary inside the same interface.
One request may receive deeper reasoning. Another may take the fast path. A smaller fallback may appear after a usage threshold. The user sees a continuous product while the machinery underneath changes.
Manually selected reasoning modes are less exposed. ChatGPT Thinking or Pro, Claude’s higher effort settings and Gemini Pro with higher thinking settings are intended for work where depth matters. Named models used through an API provide more control because the customer can usually specify the model and reasoning effort.
Even these are not fixed appliances. Providers update models, inference systems and context handling. None gives the user a dependable meter showing how much useful thought, verification or source checking went into each response.
Grok is harder to assess because xAI publishes less detail about consumer routing and how it allocates compute. Similar economic pressure is a reasonable inference. It is not evidence of identical behaviour.
Capability is improving while the service becomes less dependable
The strongest current models are substantially better than their predecessors at coding, research, tool use and multi-stage reasoning. Calling this straightforward technological decline would miss the point.
The models can become more capable while the average service becomes less reliable. Frontier intelligence exists, but the product decides when the user receives it.
That is the beginning of AI enshittification. Attract users with unusually capable systems. Let those systems become part of their work. Then ration the expensive capability, package dependable reasoning into higher tiers and preserve one reassuring interface across several grades of assistance.
The segmentation itself is not objectionable. The opacity is.
A notice saying “fast model, limited reasoning” would tell a user how much confidence to place in the result. A warning that earlier context had been compressed would explain why established instructions were forgotten. Clear notification that a usage threshold had triggered a smaller model would prevent the user assuming continuity where none existed.
Instead, the expert and the intern wear the same name badge on their lapel.
Treat fluency as presentation, not proof
For consequential work, select the strongest reasoning model available rather than leaving the decision to Auto. Give it an instruction that makes the standard explicit:
Treat this as professional work. Do not estimate or fill gaps with plausible assumptions. Research and verify material claims, distinguish evidence from inference, and tell me what remains uncertain.
That cannot force your AI provider to allocate unlimited computing resources, and it does not turn an unsuitable AI model into an expert. It does make the expected behaviour clear and gives you something concrete against which to judge the answer.
Long projects also need periodic restatement of their controlling facts and constraints. Research should require named sources. Material decisions should be checked against the underlying evidence, preferably outside the model that produced the recommendation.
The original promise was expert-level assistance at negligible marginal cost. The service now looks more like an agency with an excellent senior team, a large intake of juniors and an invisible line manager deciding who handles each brief.
You may receive the expert. You may receive the intern. Both will submit the work on the same letterhead. For now, you can improve the odds of receiving the former. But ultimately, the responsibility for discovering which one turned up lands with you.
Sources

