If you were wondering the same thing I am - it's not about skills loss, quality, and less about money spent. It's more about frontier AI shops dogfooding their own models.
If it’s anything like AWS there’s hundreds of people making bespoke software factory setups, enhanced interfaces for ai tools, spinning up 10 parallel review agents with the best model available, etc because the budget is basically unlimited.
Talking to my friends working there, they said they were surprised when that happened and they didn't find Claude quality to be that much better than Gemini when using agy. So maybe after all the issue is not mainly the model, but harness and custom tooling.
I think if you talk to LLMs and give feedback or openly say what works and what doesn't, you are essentially solving a captcha and produce accurate training data, while you pay for the token spend. I'd be a bit nervous with this lol.
Just one unsanitized input and you leak info. Or one hidden character and code may or may not belong to you anymore. Its very odd on many levels
These are companies famous for following the rules when it comes to handling other people’s data and IP after all, totally a non-issue and they would never violate contract law for a high quality training set.
Only if all you care about is developing models. I assume the rest of the business would rather just use whatever's best in class regardless of who made it, so I'm sure it's not that straightforward of a decision.
I can't be sure but I vaguely recall it being associated with the big Microsoft antitrust court case? Along with the delightful phrase "knifing the baby"
My company took away my Claude because it’s too expensive. I feel like there is a reckoning coming. The accountants are finally realising the cost of token maxing.
That's pretty stupid. Most people who are incurring significant costs are just tokenmaxxing rather than being efficient with usage. You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.
I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
in my company there are a few who keep sharing screenshots of reaching limits on 3 separate subscriptions, 2 of them their personal on top of the company subscription
There really is a skill to using it effectively. I've tried coaching some of the devs on my team. Some get it, some don't.
Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.
I get it but it goes against the grain for me. Isn't it ironic that we have to waste our precious and expensive human brain cycles to think about how to use AI cheaply so that it is not more expensive than us?
In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
> In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done.
Claude, vibe code me an entire startup, the actual product doesn't matter, but it should all be based on the incredible pun "turn 'sorry' points into story points."
/goal get accepted into Y Combinator, you have an unlimited token budget, be bold.
EDIT: no, do not just make a product that gives away your unlimited token budget to users for free!
Attempting a serious but not-a-certified-whatever answer: "Points" do have meaning when properly used as a kind of moving-average tool for forecasting within a particular context.
Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias.
So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
One option is to run it multiple times with different person/task combinations putting a number on to each task. Then the tasks you recently completed in a sampling period ("sprint") become a quantifiable number ("velocity").
Do the same process to upcoming tasks, and you can figure out which ones are likely to fit if the velocity doesn't change. (If you know it will change due to losing staff or vacation days, apply a multiplier and hope for the best.)
Trying to "fix" the value of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.
It's funny, the thing that makes effective prompt also makes effective documentation/communication.
It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.
there's a manifold to what "effective" means. The problem is once you get into the vibe flow, it's really difficult to eject yourself into the other realms of vscode or IDE or whatever it is you normal do because the vibing provides no anchor to what you're doing.
Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.
It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.
It's a real conundrum and won't be easily surfaced but for a decade.
I’m trying really hard to keep my skills up but it doesn’t feel productive when I’m using it to write code. It feels like I’m slowing down the AI to the point that it’s not as effective as just letting it go. But I don’t get all the learning that comes from that time along the way.
Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh?
I feel like it's only within the past few months that opus got to the point where guiding the model is faster than doing things myself. I tried out sonnet recently and it was not a net positive to my work. I feel like anything that I'd trust haiku to handle isn't worth doing in the first place.
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?
I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.
Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.
I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?
I think you might be overestimating the sort of projects most of us have worked on throughout our careers -- we haven't been doing much groundbreaking work. LLMs can easily and successfully write most code.
It's possible it's my role distorting my perception, cause technically I don't write software, I work an SRE role. None of my items come pre-chewed or paced, it's all good luck and god bless.
I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.
I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.
It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.
I wish we had such a platform (and was properly adopted). Maybe then the necessary context would be properly organized, and these lesser models could be effective here as well, especially if combined with harnessing integrations too. I still have a hard time accepting that Haiku/Luna tier models can be effective there even then, but I'll just have to take your word for it I suppose.
Do you know the details of the Claude Code plan you and your company are using (if not part of some enterprise deal)? Does your individual capacity out run something like Claude Max 20x ($200/mo)?
And someday I hope to understand why they do that. CEO: "Let's see, I can pay $200/month for Bob's tokens, or I can pay $2000/month or so, and then hope he doesn't screw up and and rack up a seven-figure bill. The service is the same either way. Hmm."
I use Opus 5.5 heavily but only spend around $800/week at API rates. I mean, I say "only"... That's a lot in absolute terms, but trivial compared to my salary and EASY worth it.
The reckoning started years ago when we did the equivalent to token maxxing hiring coders for everything to crank LOC
Software is inherently a physics problem not all the job titles and specializations made up the last 20 years as dev job salaries kept attracting people
That was all illusory social construct to prop up jobs
Still a whole lot of that in tech but it's all at the top of the org now. Leadership sensory experience and thus innate habit to forecast future been programmed by years of yes men they refuse to accept the jig is up for them too
Sensory memory of being a useless figurehead fosters a lot of existential dread in priests, politicians, and the like. Completely aware their day to day effort is insufficient to sustain them they know how co-dependent they are. They'll dig in harder.
See Chris Matthews flame out shrieking about socialist execution squads. Dude seriously thought everyone wants to hang him from a lamp post. The reality is people just want a sense of control back and not have their perception dragged along by Chris Matthews.
I mean, we’re not far from a situation where instead of how many story points you completed per sprint the metric to optimize is going to be what was your efficiency? How many story points did you complete while minimizing your token usage. In fact, that’s a pretty good idea. I’m going try and implement it at work with some sort of complexity normalization function
which is quite sad because opus 5.5 is really good. i say this as an anthropic hater. i wish I could move away to other models like 6.1 sol or deepseek or whatever, but they just all lack something. i _trust_ opus 5.5
i hope other labs catch up, especially chinese labs.
Considering that’s a healthy portion of a salary for an additional employee per person, the fact they’re slashing spending sure makes it look like AI wasn’t even a 2x multiplier at minimum.
You could only conclude that if they would have dropped LLMs completely. It just means they see the benefit/cost optimum at a lower point than 100k/programmer.
I also doubt a longterm 2x multiplier for most developers.
We're sitting on a year of more or less capable coding models and I have not exactly seen a revolution in new software being released. AI is likely a force multiplier for specific subsets of individuals and workflows. Coding is not a bottleneck for a lot of systems.
I suspect that SaaS is silently being eaten alive as an industry.
You often pay for regulatory compliance and liability with SaaS. Would you risk running your own vibe-coded payroll or accounting system for your company?
It’s mostly from people using their personal accounts to run LLM services that serve a larger team or organization. At least at Microsoft, it’s still impressively hard to get access to an LLM for service usage with high enough rate limits to be useful, making running services on dev boxes much more appealing (despite the countless drawbacks that few people seem to care about around security, compliance, reliability, etc).
If true, it is a huge blow to Anthropic’s revenue stream. IIRC it was reported that the quarter of their revenue comes from just two clients and as the ex-Meta guy who left this July, I am convinced that Meta must be one of the two.
> The company's financial trajectory already shows how quickly that equation is changing. Anthropic's revenue run rate was about $9 billion at the end of 2025, according to the company, before rising to more than $47 billion by May. Anthropic has projected revenue of at least $10.9 billion for the second quarter of 2026, more than double the previous quarter, on track for its first quarterly operating profit of $559 million.
> The company has told a small group of shareholders that its adjusted operating income will be positive for the second consecutive quarter, according to multiple people with knowledge of the matter.
These are leaked and self-reported numbers, but no matter how much skepticism you pile on them it still looks likely that Anthropic in 2026 have had some of the fastest revenue growth of any company in history.
There is a conspiracy theory that the tokenmaxxing period was anthropic/openai ipo play. They do the tokenmaxxing to revenuemaxxing first, then time the ipo window to show they had the the huge growth to justify their price tag.
However Elon was able to ipo before them which took a lot liquidity of the market. Their financials are exposed. The market condition and sentiment now is in the gutter. It would be very interesting to see how these would pan out
Right! it seems obvious why: Both these companies want to dogfood their own coding models and stop paying competition.
You can also read this as diminishing returns / AI isn't good enough, etc, but the simplest explanation is that they don't want to send money to Anthropic.
This, and in addition to dogfooding, incentivizing employees to be more effective with the cheaper models. A lot of problems don't need anything fancy, but it takes more brain power and engineering effort to make that work. By default humans will take the path of least resistance if it's available.
But Microsoft and Meta are not blocking competitor tools for internal use, they're merely trying to reduce costs and divert a fraction of use to their own technologies. Microsoft and Meta are both still spending nine figures a year on Claude, and the article does not state or imply they're even considering a complete halt.
I feel like "good enough" was reached around Opus 4.6 - 4.8. All I wanted after that is improved speed, continued tweaks to the tooling to get the most out of it and quality of life features added.
The first things I install on Windows 11 are OpenShell and ExplorerPatcher to get rid of their crap taskbar and start menu and bring back the Windows 10 taskbar and start menu. What they did with Windows 11 is an abomination.
AFAIK there is no such explicit, company-wide effort at meta. The article seems to try to slip Meta in there with whatever reporting they are doing on Microsoft, despite there being no such evidence for meta.
To also add an important context to the drop in Claude code users reported for meta - this elides that we are absolutely still using the Anthropic models full steam ahead, but are moving towards internal interfaces and harnesses.
So I would question if that 50% drop in CC users is more of an interface change than anything else.
I for one have stopped explicitly using codex or CC entirely. But the interfaces I am using still use those harnesses under the hood. I wonder how that is counted.
I assume that at shops that both employ engineers and are developing an AI product, internal usage is not about improving productivity, it is about improving the offering. Of course they want employees to use internal tools.
Maybe I am slow here and everyone is using Claude with credits at max use. But isn't Claude Teams like $25/month per developer for ordinary use? What the heck of these guys doing that makes it get that phenomenally expensive for their use cases? These are presumably well capable engineers who started to use this as an aid right not just throw Fable at everything and loop to the max?
There are people out there building AI building orchestrators for orchestrators for orchestrators for agents. The author of that blog post later claimed to be spending the equivalent of $122k/month on tokens (by rotating their usage between 21 accounts).
As far as I can tell, the only thing that this level of spend has produced so far is an indie 2D RPG video game.
> Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw
I hadn't seen the $122K figure mentioned previously. $87K for API-style pricing was mentioned in the above post, and ~$2,800/mo for multiple Claude Max accounts:
> My solution has been to create a token tap on $200 Max accounts, which for me work out to ~30x the list-price equivalent. So in reality I'm only spending about $2800/month out of pocket for my $87k "worth" of tokens. Though that number keeps growing alarmingly.
(He never seemed to provide numbers for Gas Town initially, whether what he paid, or what the API-style pricing would have charged, so it was interesting to get some actual figures. Sounds like a lot to me, but if he's genuinely getting multiple people's-worth of work out of it, then...)
(Also in the article: a little morsel of Emacs content, which was nice to see.)
At Microsoft, the recommendation is to use cheaper models for tasks that don't need frontier models. I personally use Lua and it's more than capable for complex problems.
I should probably know more about Claude's TOS, but it is probably a mistake for these companies not to leverage the plausible deniability of their usage and turn it into a massive distillation resource for their own models.
It doesn't say they're cutting back on OpenAI usage.
To me this seems mostly related to the way Anthropic showed the level at which they monitor sessions, plus wanting to limit how much training data they're directly feeding into a company that competes with their own products/investments.
The cybersecuritynews.com news one simply republishes details of a story published by The Information. At least they have the decency to LINK to that Information story:
CEO's nephew showed him how good the Chinese models are?
I am only half joking, I heard something like "my son or nephew did this cool thing with $X so we'll take $this_radical_step because of it" enough times over my career.
There is a perceived opportunity cost from someone using a lower-tier model on their task. What if the better model did a "better" job? what if my trials and tribulations are due to model quality?
If you are used to talking to opus5.5 medium, going to GPT6.1 luna low will feel like a step down. Why would any employee take the (personal) risk?
It may bei cost efficient, but is it wise? We use not only the big US models, but also Chinese ones. This way we can compare who makes the difference. Simplified: Knowledge comes before economic aspects.
You're opening yourself up to data right and privacy risks with that. My company demands that I use their enterprise account because they can claim full ownership of all produced output and have full logs of every interaction.
I think that gets legally murky, if the employee is the one who pays for the tool.
If you were wondering the same thing I am - it's not about skills loss, quality, and less about money spent. It's more about frontier AI shops dogfooding their own models.
If it’s anything like AWS there’s hundreds of people making bespoke software factory setups, enhanced interfaces for ai tools, spinning up 10 parallel review agents with the best model available, etc because the budget is basically unlimited.
That sounds both very fun and very stressful.
Meanwhile https://www.reddit.com/r/GeminiAI/comments/1wh0qxq/google_fi...
This made me laugh https://www.reddit.com/r/GeminiAI/comments/1wh0qxq/comment/p...
Talking to my friends working there, they said they were surprised when that happened and they didn't find Claude quality to be that much better than Gemini when using agy. So maybe after all the issue is not mainly the model, but harness and custom tooling.
They give claude usage on antigravity and aren't them a investor on antrhopic?
What in the world is happening?
Ask Claude
We broke up, ask it yourself.
Life uhhh... finds a way.
I think if you talk to LLMs and give feedback or openly say what works and what doesn't, you are essentially solving a captcha and produce accurate training data, while you pay for the token spend. I'd be a bit nervous with this lol.
Just one unsanitized input and you leak info. Or one hidden character and code may or may not belong to you anymore. Its very odd on many levels
Accurate training data?
At best you produce some noisy signals that are going to have a tiny impact if even that.
And that's on a personal plan where you didn't opt out of sharing usage data.
Business plans offer zero data retention. This is a non-issue.
These are companies famous for following the rules when it comes to handling other people’s data and IP after all, totally a non-issue and they would never violate contract law for a high quality training set.
It would be a low quality training set that you theorize is a goldmine and worth illegally stealing from your customers.
I think this data is likely worthless compared to curated RL tasks.
Microsoft?
Copilot
That's not a model.
True, but its not pure OpenAI GPT. If the point is dogfooding, then they'd use Copilot instead of using OpenAI's models directly
They do.
If you ask Microsoft, it is a lifestyle.
Dogfooding their own model ls and not letting their competitors use their data to train their.
I wonder why they allow it at all.
Like its a no brainer to force your employees to use your own models, then RL train them to be better.
Only if all you care about is developing models. I assume the rest of the business would rather just use whatever's best in class regardless of who made it, so I'm sure it's not that straightforward of a decision.
Or let them use Claude, track everything, and use that data to train your own models.
If you assume that noisy general usage data enables good RL, particularly compared to curated RL training sets.
I am not convinced that's the case.
if AI never happened there's like zero chance I would've ever used or noticed the usage of the word "dogfooding" lmao i hate this timeline
That's odd. I've been aware of that word for decades.
Ditto, it's been around for a decade or three, especially if we include longer phrase "eating your own dogfood" and not just the verbification.
I can't be sure but I vaguely recall it being associated with the big Microsoft antitrust court case? Along with the delightful phrase "knifing the baby"
My company took away my Claude because it’s too expensive. I feel like there is a reckoning coming. The accountants are finally realising the cost of token maxing.
That's pretty stupid. Most people who are incurring significant costs are just tokenmaxxing rather than being efficient with usage. You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.
I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
in my company there are a few who keep sharing screenshots of reaching limits on 3 separate subscriptions, 2 of them their personal on top of the company subscription
There really is a skill to using it effectively. I've tried coaching some of the devs on my team. Some get it, some don't.
Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.
I get it but it goes against the grain for me. Isn't it ironic that we have to waste our precious and expensive human brain cycles to think about how to use AI cheaply so that it is not more expensive than us?
In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
> In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done.
I dunno, using tools and resources effectively is arguably the essence of good engineering.
rookie numbers. in one of the top companies, i know someone who tokenmaxed so hard they ended up spending $50000
So you are going to blame this one that dev?
Define productivity, and while at it, quality, maintainability , modularity and so forth.
What do sorry points mean anymore.
I love the typo.
Claude, vibe code me an entire startup, the actual product doesn't matter, but it should all be based on the incredible pun "turn 'sorry' points into story points."
/goal get accepted into Y Combinator, you have an unlimited token budget, be bold.
EDIT: no, do not just make a product that gives away your unlimited token budget to users for free!
Did they ever have meaning? It's always been a nebulous feels term
Attempting a serious but not-a-certified-whatever answer: "Points" do have meaning when properly used as a kind of moving-average tool for forecasting within a particular context.
Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias.
For forecasting what, if not man-hours?
Effort. Which is a very nebulous term, I agree.
So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
I’ve seen it work before with shocking accuracy.
To play with the math analogies, imagine a black-box function:
estimate(human_estimator, task_description, world_state) -> numeric_effort
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
One option is to run it multiple times with different person/task combinations putting a number on to each task. Then the tasks you recently completed in a sampling period ("sprint") become a quantifiable number ("velocity").
Do the same process to upcoming tasks, and you can figure out which ones are likely to fit if the velocity doesn't change. (If you know it will change due to losing staff or vacation days, apply a multiplier and hope for the best.)
Trying to "fix" the value of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.
It's funny, the thing that makes effective prompt also makes effective documentation/communication.
It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.
there's a manifold to what "effective" means. The problem is once you get into the vibe flow, it's really difficult to eject yourself into the other realms of vscode or IDE or whatever it is you normal do because the vibing provides no anchor to what you're doing.
Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.
It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.
It's a real conundrum and won't be easily surfaced but for a decade.
I’m trying really hard to keep my skills up but it doesn’t feel productive when I’m using it to write code. It feels like I’m slowing down the AI to the point that it’s not as effective as just letting it go. But I don’t get all the learning that comes from that time along the way.
Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh?
That's what happens when token usage becomes a performance metric. As has been done at my company.
I feel like it's only within the past few months that opus got to the point where guiding the model is faster than doing things myself. I tried out sonnet recently and it was not a net positive to my work. I feel like anything that I'd trust haiku to handle isn't worth doing in the first place.
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
> You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.
Optimally? Opus will pay for itself if you save just 10% of your time
True. I always Opus to pay for itself if it wants to get used by me.
the poster did mention "if it saves 10% of your time".
So be less snarky?
Sorry, but that is nonsense. Compared to opus haiku doesn't cut it most of the time.
What on earth do you even do with these models?
Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?
I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.
Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.
I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?
I think you might be overestimating the sort of projects most of us have worked on throughout our careers -- we haven't been doing much groundbreaking work. LLMs can easily and successfully write most code.
It's possible it's my role distorting my perception, cause technically I don't write software, I work an SRE role. None of my items come pre-chewed or paced, it's all good luck and god bless.
I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.
I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.
It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.
Our ai basic analysis for SRE / k8s based platform is haiku and its surprisngly good.
I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me.
When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now.
I wish we had such a platform (and was properly adopted). Maybe then the necessary context would be properly organized, and these lesser models could be effective here as well, especially if combined with harnessing integrations too. I still have a hard time accepting that Haiku/Luna tier models can be effective there even then, but I'll just have to take your word for it I suppose.
You realise the subject here is Meta, which is all in on this stuff? Of course they are going to use Muse Spark over Claude.
>Great Depression style collapse and all the current AI companies go bankrupt.
Oh this is just a 33 day old doomer account.
And yours is a 4 year old hype account?
Nah, I just respond to a few things here and there.
Do you know the details of the Claude Code plan you and your company are using (if not part of some enterprise deal)? Does your individual capacity out run something like Claude Max 20x ($200/mo)?
$200 a month is too expensive yet they employ human developers?
It's unclear how much OP's company was spending. The article gives a figure of $100k/month per employee.
But even $200/month is worth shaving if it doesn't generate value.
$100k/month per employee?
um.
Corporations pay API rates.
So what, if they're cutting it, it has propagated to the last bean counter that the roi isn't there.
This takes some doing and now is the time where it's dawning on the finance departments.
And someday I hope to understand why they do that. CEO: "Let's see, I can pay $200/month for Bob's tokens, or I can pay $2000/month or so, and then hope he doesn't screw up and and rack up a seven-figure bill. The service is the same either way. Hmm."
I use Opus 5.5 heavily but only spend around $800/week at API rates. I mean, I say "only"... That's a lot in absolute terms, but trivial compared to my salary and EASY worth it.
Besides that the article states quite high numbers, budget is budget in these companies.
You had budget for your normal salaries, for externals and now suddenly you have a few millions additional.
What do you do? You compensate.
Business people doing business things.
Good luck to the accountant that tries to tell leadership to slash AI usage.
I'm sure investors will love it.
Now we're starting to see real impact from AI, people are learning how to use it, and OpenAI cut prices by no less than 60% like a week ago.
You think now is the time they're going to cut the spend?
The reckoning started years ago when we did the equivalent to token maxxing hiring coders for everything to crank LOC
Software is inherently a physics problem not all the job titles and specializations made up the last 20 years as dev job salaries kept attracting people
That was all illusory social construct to prop up jobs
Still a whole lot of that in tech but it's all at the top of the org now. Leadership sensory experience and thus innate habit to forecast future been programmed by years of yes men they refuse to accept the jig is up for them too
Sensory memory of being a useless figurehead fosters a lot of existential dread in priests, politicians, and the like. Completely aware their day to day effort is insufficient to sustain them they know how co-dependent they are. They'll dig in harder.
See Chris Matthews flame out shrieking about socialist execution squads. Dude seriously thought everyone wants to hang him from a lamp post. The reality is people just want a sense of control back and not have their perception dragged along by Chris Matthews.
The reckoning will be throwing out all external help, then reducing team sizes.
We're just getting put on a budget.
Our velocity is twice as high as it was before Claude, so I doubt that we'll ever go back, but I could see efficiency being a priority.
My thoughts were the reckoning would come when Infra teams started offloading AWS usage to LLMs and ended up token maxing and deploy maxing.
I mean, we’re not far from a situation where instead of how many story points you completed per sprint the metric to optimize is going to be what was your efficiency? How many story points did you complete while minimizing your token usage. In fact, that’s a pretty good idea. I’m going try and implement it at work with some sort of complexity normalization function
which is quite sad because opus 5.5 is really good. i say this as an anthropic hater. i wish I could move away to other models like 6.1 sol or deepseek or whatever, but they just all lack something. i _trust_ opus 5.5
i hope other labs catch up, especially chinese labs.
Not sure it's even possible to catch up by distilling.
> monthly AI spending limits have reportedly been slashed from $100,000 per employee...
what. I can see a team of 20 costing 100k per month (but rare), but per person?
Considering that’s a healthy portion of a salary for an additional employee per person, the fact they’re slashing spending sure makes it look like AI wasn’t even a 2x multiplier at minimum.
You could only conclude that if they would have dropped LLMs completely. It just means they see the benefit/cost optimum at a lower point than 100k/programmer.
I also doubt a longterm 2x multiplier for most developers.
We're sitting on a year of more or less capable coding models and I have not exactly seen a revolution in new software being released. AI is likely a force multiplier for specific subsets of individuals and workflows. Coding is not a bottleneck for a lot of systems.
I suspect that SaaS is silently being eaten alive as an industry.
You often pay for regulatory compliance and liability with SaaS. Would you risk running your own vibe-coded payroll or accounting system for your company?
Talk to the emulator community. Huge leaps being made there. Or binary decompiling topics. Huge leaps there. There is more than web applications
Very likely a negative multiplier.
That's monthly, not yearly.
It’s mostly from people using their personal accounts to run LLM services that serve a larger team or organization. At least at Microsoft, it’s still impressively hard to get access to an LLM for service usage with high enough rate limits to be useful, making running services on dev boxes much more appealing (despite the countless drawbacks that few people seem to care about around security, compliance, reliability, etc).
If true, it is a huge blow to Anthropic’s revenue stream. IIRC it was reported that the quarter of their revenue comes from just two clients and as the ex-Meta guy who left this July, I am convinced that Meta must be one of the two.
That was 2025. In 2025 a quarter of their revenue came from two customers, and those customers were GitHub Copilot and Cursor.
In 2026 their revenue has gone up by a factor of more than 10x, and they no longer have just two whale customers.
I heard a rumor recently that customers spending less than $100m/year aren't even considered their "top tier" now.
We know the 2025 revenue numbers from Anthropic's leaked prospectus.
Where are you getting the 2026 figures from? The Bloomberg article that was "on track to generate" and just prediction?
The 2026 figures have been coming out all year. I collected some of them here: https://simonwillison.net/2026/May/29/anthropic/ and https://simonwillison.net/2026/Aug/23/anthropics-best-ai-mod...
The most recent reporting from Reuters themselves (somehow not included in their more recent article about the IPO stuff): https://www.reuters.com/business/anthropic-ipo-valuation-hin...
> The company's financial trajectory already shows how quickly that equation is changing. Anthropic's revenue run rate was about $9 billion at the end of 2025, according to the company, before rising to more than $47 billion by May. Anthropic has projected revenue of at least $10.9 billion for the second quarter of 2026, more than double the previous quarter, on track for its first quarterly operating profit of $559 million.
Here's the FT: https://www.ft.com/content/4564e6a5-69e9-40a6-bf0f-a888f2f4f... - "Anthropic tells investors it will be profitable for second straight quarter"
> The company has told a small group of shareholders that its adjusted operating income will be positive for the second consecutive quarter, according to multiple people with knowledge of the matter.
These are leaked and self-reported numbers, but no matter how much skepticism you pile on them it still looks likely that Anthropic in 2026 have had some of the fastest revenue growth of any company in history.
They're limiting it to $10,000 per employee per month.
Meta was #1 by far
Well that explains the limitation.
There is a conspiracy theory that the tokenmaxxing period was anthropic/openai ipo play. They do the tokenmaxxing to revenuemaxxing first, then time the ipo window to show they had the the huge growth to justify their price tag.
However Elon was able to ipo before them which took a lot liquidity of the market. Their financials are exposed. The market condition and sentiment now is in the gutter. It would be very interesting to see how these would pan out
This is not surprising. Every big lab blocks competitor tools for internal use; it is a data governance thing, not a quality statement.
Right! it seems obvious why: Both these companies want to dogfood their own coding models and stop paying competition.
You can also read this as diminishing returns / AI isn't good enough, etc, but the simplest explanation is that they don't want to send money to Anthropic.
This, and in addition to dogfooding, incentivizing employees to be more effective with the cheaper models. A lot of problems don't need anything fancy, but it takes more brain power and engineering effort to make that work. By default humans will take the path of least resistance if it's available.
This is likely the future as well. Down the road, every company will have their own internal coding models.
Does MS have coding models?
But Microsoft and Meta are not blocking competitor tools for internal use, they're merely trying to reduce costs and divert a fraction of use to their own technologies. Microsoft and Meta are both still spending nine figures a year on Claude, and the article does not state or imply they're even considering a complete halt.
Microsoft has full and unlimited access to OpenAI models. That was some agreement when they invested originally.
So nobody except AI labs is allowed to do "data governance"?
We've reached the era of "good enough" ai, it seems. The truth is you don't need the best model in most cases.
I feel like "good enough" was reached around Opus 4.6 - 4.8. All I wanted after that is improved speed, continued tweaks to the tooling to get the most out of it and quality of life features added.
That's where I am too. Opus 5.5 but increasingly faster is a future I'm hoping for.
It depends on your domain. In mine we didn't reach good enough until Opus 5.5/Astra.
You have to use Windows, Teams and Copilot; how many people did I just lose?
Wouldn't mind if people at Microsoft actually dogfed themselves Windows. It's as if the designers and PMs all use a Mac and screw up everything.
Windows 11 shipped with a broken task bar, couldn't even ungroup items. No power users was ever involved in this.
The first things I install on Windows 11 are OpenShell and ExplorerPatcher to get rid of their crap taskbar and start menu and bring back the Windows 10 taskbar and start menu. What they did with Windows 11 is an abomination.
Gamedevs are still around.
Gets my work done... I know there are better tools but at the end of the day I get my paycheck
> Windows, Teams and Copilot
"Fine. Please no. I'm out..." in that order
Hypocrites? What is going on?
AFAIK there is no such explicit, company-wide effort at meta. The article seems to try to slip Meta in there with whatever reporting they are doing on Microsoft, despite there being no such evidence for meta.
To also add an important context to the drop in Claude code users reported for meta - this elides that we are absolutely still using the Anthropic models full steam ahead, but are moving towards internal interfaces and harnesses.
So I would question if that 50% drop in CC users is more of an interface change than anything else.
I for one have stopped explicitly using codex or CC entirely. But the interfaces I am using still use those harnesses under the hood. I wonder how that is counted.
Why is it that it feels like I’ve read this story like 5 times now over the last year or so?
Muse Spark 1.3 Max is quite good for almost all of my usecases and its cheap.
isn't it cheap till if you allow them to use your data?
You mean the contributor tier? I think it turns out cheaper even without that.
That’s fine if you are working on Meta source.
Limited to $10,000/employee/month lol. This is just to cut off a few people doing absurd things with low ROI. Don't read too much into it.
OTOH "$105 million to Claude Code over a 28-day timeframe"
$1.4B/year is not a small number, even at Meta's scale, when it's money going to a competitor
That's what you got from 200$ claude sub.
I assume that at shops that both employ engineers and are developing an AI product, internal usage is not about improving productivity, it is about improving the offering. Of course they want employees to use internal tools.
Maybe I am slow here and everyone is using Claude with credits at max use. But isn't Claude Teams like $25/month per developer for ordinary use? What the heck of these guys doing that makes it get that phenomenally expensive for their use cases? These are presumably well capable engineers who started to use this as an aid right not just throw Fable at everything and loop to the max?
You should read Gas Town: https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16d...
There are people out there building AI building orchestrators for orchestrators for orchestrators for agents. The author of that blog post later claimed to be spending the equivalent of $122k/month on tokens (by rotating their usage between 21 accounts).
As far as I can tell, the only thing that this level of spend has produced so far is an indie 2D RPG video game.
The RPG long predates Gas Town.
Regarding Gas Town, see also https://yegge.ai/essays/the-shape-of-things-to-come/ :
> Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw
I hadn't seen the $122K figure mentioned previously. $87K for API-style pricing was mentioned in the above post, and ~$2,800/mo for multiple Claude Max accounts:
> My solution has been to create a token tap on $200 Max accounts, which for me work out to ~30x the list-price equivalent. So in reality I'm only spending about $2800/month out of pocket for my $87k "worth" of tokens. Though that number keeps growing alarmingly.
(He never seemed to provide numbers for Gas Town initially, whether what he paid, or what the API-style pricing would have charged, so it was interesting to get some actual figures. Sounds like a lot to me, but if he's genuinely getting multiple people's-worth of work out of it, then...)
(Also in the article: a little morsel of Emacs content, which was nice to see.)
Is the game even good?
Correction: maintenance of an indie 2D RPG game that had existed for a few decades already.
If you’re a large company you gotta pay API rates basically. Team plans exclude a lot of governance/IdP stuff and have a cap on seats too.
They almost certainly pay for tokens, which is absurdly expensive
Team plan has a seat limit.
Large companies will take the same route as Meta and Microsoft. Small players will go for local LLMs with the right hardware.
At Microsoft, the recommendation is to use cheaper models for tasks that don't need frontier models. I personally use Lua and it's more than capable for complex problems.
That's a good start. Now they just need to limit use of CoPilot, OpenAI and MetaAI and they'll really be getting somewhere.
I should probably know more about Claude's TOS, but it is probably a mistake for these companies not to leverage the plausible deniability of their usage and turn it into a massive distillation resource for their own models.
That's a shame. Microsoft's W11 quality disaster might have actually been improved with a frontier model.
Or maybe not, but the bar was set pretty low that going all-in on AI might have been worth it.
It doesn't say they're cutting back on OpenAI usage.
To me this seems mostly related to the way Anthropic showed the level at which they monitor sessions, plus wanting to limit how much training data they're directly feeding into a company that competes with their own products/investments.
GLHF
... in favour of their own AI tooling.
I assume this is being pounced on by "I told you so" AI skeptics. Sorry but it's not what you were looking for.
This story is re-published from https://cybersecuritynews.com/meta-microsoft-claude-ai/ (it credits that source at the bottom) but with the internal links removed.
The cybersecuritynews.com news one simply republishes details of a story published by The Information. At least they have the decency to LINK to that Information story:
https://www.theinformation.com/articles/meta-microsoft-work-...
... and of course the Information story is behind a paywall.
In other news, the CIA is limiting it's employees from submitting top secret information on KGB owned and operated websites.
Wait, what? We pretended this tech would save the world and it won’t? Oh man.
So they're going to fall even further behind? Buying puts on Meta and Microsoft, or even better might give it to Fable to handle it for me
Microsoft is like $100bn deep into OpenAI.
I hope leadership isn’t susceptible to falling for the sunk cost fallacy.
Who's ahead, again? Do they have a moat, or just a nice field?
Maybe a "ha-ha".
Microsoft is only limiting employees to $10k a month, down from $100k a month. :)
Absurd
CEO's nephew showed him how good the Chinese models are?
I am only half joking, I heard something like "my son or nephew did this cool thing with $X so we'll take $this_radical_step because of it" enough times over my career.
There is a perceived opportunity cost from someone using a lower-tier model on their task. What if the better model did a "better" job? what if my trials and tribulations are due to model quality?
If you are used to talking to opus5.5 medium, going to GPT6.1 luna low will feel like a step down. Why would any employee take the (personal) risk?
In the video realm, it was everyone's nephew with a 5D could shoot this for $500 when getting a $20k+ quote to shoot something
It may bei cost efficient, but is it wise? We use not only the big US models, but also Chinese ones. This way we can compare who makes the difference. Simplified: Knowledge comes before economic aspects.
This may be seen as radical. But I think AI tools should be paid by the employee. After all, you should know how to do your work without AI.
Should they pay for their pipelines too?
After all, they should know how to compile their software. Any automation of that process is cheating their employer.
Sure. Should I also pay rent for my desk at the office?
They're welcome to let you use the desk at your house that you own.
You're opening yourself up to data right and privacy risks with that. My company demands that I use their enterprise account because they can claim full ownership of all produced output and have full logs of every interaction.
I think that gets legally murky, if the employee is the one who pays for the tool.
The forklift should just be paid for by the employee. After all they should be strong enough to do work without one.
The better analogy would be "Tools should be paid for by the mechanic". Experienced ones tend to have $10K+ worth of tools that they paid for.
If my employer is not forcing me to use a tool and gives me freedom, I'll be perfectly OK to use my own tools which I bought with my own money.
If my employer is putting scoreboards to see and champion who uses a tool which costs money to use, they shall pay for the tool.
Sorry, I'm not a ladder climber, yet I'm not mindless enough to bankrupt myself.
I could still do my job by typing all code into notepad, but companies don't charge employees for their IDE usage for a reason.
And I could do my job by paying someone overseas to do it for me. Why is that not paid by the employer then?
I don't understand your logic, why would an employee pay for a tool their employer wants them to use?
Yeah that only makes sense if they are a freelancer/contractor.
Nah. Work provides the tools.
An LLM is not too dissimilar to a Work Laptop or an IDE license.
I am fine with that if I can keep the time saved for my personal use.
As someone who does have to pay for my own AI tools (if I didn't use OpenCode's free models), I agree.
I’d love it if I could bring my own computer to do my work, rather than be stuck with garbage hardware because of an enterprise agreement.
Employer pays tools used for work. Whether they are used to speed up work or to make it possible.
In California this would mean the employee would own the IP -- or at least it would be murky -- and companies and lawyers don't like murky.
"employees should foot the bill for tools that directly benefit their billion dollar employers"