Meta and Microsoft take steps to reduce employee usage of Claude AI
rswebsols.com
[118 comments hidden]
[3 comments hidden]
i hope other labs catch up, especially chinese labs.
[67 comments hidden]
I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
[32 comments hidden]
Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.
[14 comments hidden]
[10 comments hidden]
[9 comments hidden]
Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias.
[5 comments hidden]
[4 comments hidden]
So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
I’ve seen it work before with shocking accuracy.
[hidden]
estimate(human_estimator, task_description, world_state) -> numeric_effort
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
A popular option is to run it multiple times with different person/task combinations, putting a projected number on to each task. Afterwards, the tasks finished in sampling period ("sprint") become a quantifiable total for that period ("velocity").
Do the same process again with the next set of tasks, and you can figure out which ones are likely to fit if the velocity doesn't change much. If you know the velocity will change due to losing staff or vacation days... well, we apply a multiplier and hope for the best.
Trying to "fix" the meaning of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.
[hidden]
[3 comments hidden]
My wife, despite loving all things French, just doesn't "do metric". She wants my height in feet and inches, my weight in pounds, boom done. So while I know my mass in kilograms, she needs the conversion done before she can even begin to have a reference point.
Upper management is the same way. They have forecasts that they need to make, deadlines and budget goals that they need to hit. They only deal in the units of hours and dollars (or local currency). Every software engineer is accountable for their work in those units only. The conversion needs to be done before the management chain has a reference point.
One easy way to do this is to have each engineer estimate the time it takes to fulfill a story after it's been pointed; then, upon completion, record their actual hours spent. Their estimated vs. actuals tend to stabilize over time, so even if they misestimate a task, you can arrive at a good guess at the time it will actually take.
[hidden]
> Upper management [...] only deal in the units of hours
I feel you're mixing up different operations here. You can always express unfinished work as likely to require a certain number of team-sprints, which are convertible to theoretic man-hours. The key is that the conversation rate is only valid for a moment, and technically that moment was the prior sprint.
That's very different from management thinking (or worse, declaring) that points have a permanently fixed proportion to man-hours.
> [...] and dollars
If your management deals in international currency, then perhaps that would be a useful analogy to them: Points and Man-Hours are different sides of FOREX, and they fluctuate based on different conditions.
When Engineering predicts a group of tasks is 54 points, that's like a foreign company signing a long-term contract in €100 EUR instead of USD. You can estimate that you'll receive ~$112 USD in a year, but the actual dollars will likely be different because the exchange-rate will continue changing before that happens.
[3 comments hidden]
[2 comments hidden]
/goal get accepted into Y Combinator, you have an unlimited token budget, be bold.
EDIT: no, do not just make a product that gives away your unlimited token budget to users for free!
[3 comments hidden]
Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.
It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.
It's a real conundrum and won't be easily surfaced but for a decade.
[2 comments hidden]
Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh?
[hidden]
Define productivity, and while at it, quality, maintainability , modularity and so forth.
[6 comments hidden]
In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
[hidden]
It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done.
[hidden]
[2 comments hidden]
Nothing ironic about it, it's basic engineering efficiency optimization.
Just today I was calculating that if money was not a factor at all, we could simply use Mythos to handle all of our continuous security scanning needs, at a cost of about $15 million a year.
Well, my budget is far (far, far, far, far) below 15M a year, so that's not going to work. So I must find compromises to make it work within budget. The single most important engineering constraint is always budget. Everything would be easier with infinite money, but there is never infinite money.
[hidden]
[3 comments hidden]
It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.
[hidden]
[hidden]
[hidden]
[hidden]
Very much this. There is a vast range of effectiveness and combined with so many models and pricing tiers, it can be tricky for some. You really need to treat it like partly a programming language, but also partly as a management delegation exercise (do I delegate this task to the intern (cheapest model) or to the principal engineer (frontier model) based on complexity).
I do coworking sessions with most in my team to see how they are using it to understand and coach for effectiveness. Just using the most expensive model for everything isn't going to cut it anymore in the post-tokenmaxxing age.
[4 comments hidden]
Optimally? Opus will pay for itself if you save just 10% of your time
[7 comments hidden]
Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?
I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.
Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.
I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?
[6 comments hidden]
[4 comments hidden]
I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.
I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.
It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.
[2 comments hidden]
I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me.
When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now.
[hidden]
[hidden]
[5 comments hidden]
[4 comments hidden]
[3 comments hidden]
[2 comments hidden]
Though it's significantly slower in Token/s and also thinks a lot more without matching the same intelligence (xhigh qwen27b scores lower than haiku's medium setting, and haiku-med is $0.05 per task compared to Qwen27B's $1.01 on AA's comparison)
It seems like more a backup if you need to work offline, imo, unless time doesn't matter and/or your electricity is free. Or you want independence from the labs (fair enough).
[11 comments hidden]
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
[9 comments hidden]
There is no way I can beat even local models at generating complex Python scripts fast.
hn is filled with uber geniuses.
[hidden]
[7 comments hidden]
I had Opus trying to simplify a query for me which was slow - it ran for maybe 30 minutes, including writing and running tests, and came up with a refactor across 9 files with a couple hundred lines changed. I was looking through the output before moving onto the next step, and noticed something a little fishy- I said “why does it do x, isn’t that a more complex y?”
Opus thought for another 20-30 seconds then output “Actually that would make the majority of the diff irrelevant, if we do that change it is just these 4 lines in this single file instead.
So then I had it do that. 5-10 minutes of writing and testing and that was done.
So my company spent $25 in tokens and I spent probably an hour in total for a 4 line change that, in the days before Claude, I probably could have found the correct file and thought through the problem, understood the solution, and written the 4 lines of code myself. Probably in the same amount of time.
So basically there was no benefit at all for my time, an extra cost to the company of $25, and now I understand our codebase a little bit less instead of more if I had done all the work.
As good as Claude is at building greenfield projects it still struggles a lot at complex ones
[2 comments hidden]
And don’t forget the company also spent a bunch of money in tokens for the initial author to implement the thing poorly.
[3 comments hidden]
Dev + AI spend 3-4 hours on a project plan, there's a "wait a minute" moment, and finally they spend another hour dialing it back to a solution that could have been built, tested, and deployed in 2 hours.
Example: Someone was setting up a dev environment with multiple DB migrations from different branches - AI planned this wild 8 phase solution with a pretty fancy cutover event.
In review I essentially said... "Wait, isn't this a dev environment? It doesn't need 0 downtime, why not just destroy and recreate the DB" and it turned into a <1000LOC script.
Technically the original plan would have worked, it would have been more robust, but it would have taken a good deal more time to implement.
Some of this falls on the devs to know what fits our team well, what's realistic, what's obviously overengineered, etc... But some of it feels like AI just defaults to the most complex version of a thing. I catch it SUPER frequently. (And unfortunately some devs think that more complexity means it's a better solution)
[2 comments hidden]
I feel like I’m going crazy, using all of the best models, spending time to have excellent prompts, configuring tools and skills… and still getting overly complex solutions with mediocre results.
Like it’s still impressive how far we’ve come, and undoubtably cool technology. It’s made a bunch of personal projects possible that I never would’ve started.
But for a business I’m struggling to see the ROI. Sometimes there’s a big benefit and sometimes it’s net negative. Not saying we won’t get there but I’m trying to stay grounded in the reality of today rather than the hopes of where the technology could get to
[hidden]
But, I fear that the "facade of complexity" makes the output SEEM better. Someone might read a 10 page Codex generated plan with fancy diagrams / charts and have a feeling that because it is so complex, it must be good!
Same deal with text output in general. Oh, it uses a lot of big words and there's a LOT here, it must have done a lot of work to get to that point.
I find, in reality, that it takes much more effort to get to the simplest solution.
As they say, any old fella can build a bridge that stands, but it takes an engineer to build a bridge that barely stands...
[hidden]
The trick is making the workload manageable by the cheapest models, or costs will destroy you. We seek to make all repetitive tasks be effectively done by cursor composer model which is the cheapest.
Doing recurring tasks with anything more expensive than that will burn the budget in no time.
[hidden]
[2 comments hidden]
[2 comments hidden]
You can get the same jobs and work done with Kimi and GLM (ZDR on OpenRouter) for a fraction of the price too.
[hidden]
I'm just working in DevOps though, so it's writing IaC, not application code (save the odd Lambda function or python script). Still, even when I spend an entire day conversing with Claude and watching "bot go brrrrr", I'm one of the lowest users in our company. I have no idea what the devs who regularly hit their limits are doing.
[3 comments hidden]
>Great Depression style collapse and all the current AI companies go bankrupt.
Oh this is just a 33 day old doomer account.
[hidden]
[10 comments hidden]
Our velocity is twice as high as it was before Claude, so I doubt that we'll ever go back, but I could see efficiency being a priority.
[4 comments hidden]
Is this the new buzzword for the quarter? Last quarter was "granularity", I didn't get the memo yet
[3 comments hidden]
[2 comments hidden]
Might as well go back to counting the number of lines or number of commits. It doesn't sound as good as "granularity" and "velocity" though. The good thing with velocity is that it doesn't care which way you're going as long as you're going there fast, so you can never be wrong
[hidden]
Velocity famously being a vector with a direction. Go in the wrong direction you'd have negative velocity.
Your description would be better for the concept of speed (IE the magnitude of velocity or directionless velocity).
[5 comments hidden]
[4 comments hidden]
[3 comments hidden]
With agents I've been able to make a decent sized dent in it. Lots of this gain is due to AI - but the fact of the matter is we still have 50+ bigger projects we could work on, and non-developers building with AI has increased that number.
Busier than ever, because of AI.
[2 comments hidden]
Certainly more change is getting done, but is the market paying more for it?
[hidden]
Some of our AI tooling IS “generating revenue”, but not most of it.
Much of it is workflow efficiencies and general improvements.
One of our big initiatives is cutting out the CRM we use. It’s going to save about 30K a year to do that in house.
Our AI bill is going to be sub 30K USD this year, and if you asked the CEO if there was a return on that investment I believe they would be positive.
But, I am 100% sure there are multiple companies wasting money on AI and dialing it back because they are NOT getting a return.
[18 comments hidden]
[13 comments hidden]
[5 comments hidden]
This takes some doing and now is the time where it's dawning on the finance departments.
[4 comments hidden]
So it's not $200 a month but it can easily reach $200 a day, and unless you're a startup playing with monopoly money the maths don't work
[3 comments hidden]
A simple task, migrating a typical password reset flow as part of updating a long-lived web application from legacy libraries and software architecture to modern equivalents, apparently cost roughly the same as 2 months of Pro subscription, over the equivalent of about half a working day in wall time.
It produced code of decent quality at a small scale, but it wasn’t always on point architecturally. It also had a tendency to drift off topic and try to tangle up other changes it decided should be made with the main change we were supposed to be working towards. So even for a routine task, based on a plan developed using the harness first and with the agents working under close supervision, a near-SOTA model is still producing quality on par with a decent mid-level developer but substandard for anyone senior+ in this case.
Moreover, based on a direct comparison with other migration tasks of similar complexity that I’d already done by hand, it was actually a bit slower overall to work this way. I had to babysit Claude throughout and review everything it proposed carefully, both to avoid subtle errors (it would have made several) and to prevent drifting off track. I also had to spend a significant amount of time cleaning up its final output to an acceptable standard after the session. Those two overheads more than cancelled out the much faster code generation an LLM offers under favourable conditions.
So for now, I remain sceptical about these high multiples of improved productivity that I keep seeing claimed online from people who are apparently writing almost everything using AIs now. I could certainly have achieved a multiple of my normal productivity by YOLOing everything without reviewing it in detail and then accepting the output code without tidying anything up. However, I doubt this codebase would still have been good enough for normal human developers to work on it reasonably after even 10 or 20 AI-led sessions like that. The architecture would have degraded significantly and the test suite would have been large and largely pointless. And again, this wasn’t rocket science in this experiment, it was completely unremarkable maintenance of a relatively small and simple web application.
[hidden]
And now if you get the task done faster with the same quality. And the cost is 80 cents, considering DeepSeek and your own server starts to starts to be relevant...
[hidden]
[hidden]
[6 comments hidden]
[5 comments hidden]
[4 comments hidden]
Even if so, it might be worth spinning up a separate LLC for each division, given what they are charging for the API.
[hidden]
[hidden]
Software is inherently a physics problem not all the job titles and specializations made up the last 20 years as dev job salaries kept attracting people
That was all illusory social construct to prop up jobs
Still a whole lot of that in tech but it's all at the top of the org now. Leadership sensory experience and thus innate habit to forecast future been programmed by years of yes men they refuse to accept the jig is up for them too
Sensory memory of being a useless figurehead fosters a lot of existential dread in priests, politicians, and the like. Completely aware their day to day effort is insufficient to sustain them they know how co-dependent they are. They'll dig in harder.
See Chris Matthews flame out shrieking about socialist execution squads. Dude seriously thought everyone wants to hang him from a lamp post. The reality is people just want a sense of control back and not have their perception dragged along by Chris Matthews.
[hidden]
[hidden]
I'm sure investors will love it.
Now we're starting to see real impact from AI, people are learning how to use it, and OpenAI cut prices by no less than 60% like a week ago.
You think now is the time they're going to cut the spend?
[2 comments hidden]
[3 comments hidden]
Since then I've had fable cranked up to 11 for even the most trivial of tasks.
[2 comments hidden]
[hidden]
Disclosure: I work on Xpoze.
[hidden]
Since this summer coding on Opencode Go + Codex for a total 28$/month gives me more intelligence and token than 400$ did in may.
Also, SOTA models are increasingly useless for anything even barely tangential to security work.
[hidden]
My company didn't took away any expensive model, but it did imposed a tight budget and gave my team hardware to run local models. The budget works mainly as an influence on the decision process of which model to run. We can still run expensive budget-burning models if we really want to, but there is a clear incentive to use options that allow us to save our token budget for a rainy prompt.
Surprisingly, I feel this had a positive impact on results. Instead of succumbing to extremely expensive one-shot prompts with the most expensive model in the news, chaining agents running cheap models equiped with context and specialized skills and scripts ends up having a better and more reliable output. And faster too.
I think there is a lot of propaganda, perhaps even astroturfing, on how only the most expensive models churned out by US companies are worth using. Nowadays the cheapest models get the job done, and local models can already handle most tasks as well specially as part of orchestration chains.
[18 comments hidden]
They were allowing $100,000 per month per employee?!
That's incredible. Just a few years ago that would have been zero.
[11 comments hidden]
[5 comments hidden]
[4 comments hidden]
[2 comments hidden]
We're working on bringing the last 10% down by giving non-developers various AI skills such as a compliance module that can be pushed to specific Entra groups which will make sure their vibe code madness will stick to external dependencies we've approved. (There are still hard checks later).
For other tasks it's running on guesses, and sometimes you can save help people €1000 a week by teaching them how to use the AI better.
[3 comments hidden]
The question is does it scale properly? Answer: it depends.
What is the value of a project going from six months to six hours? You can look at saving the man hours but how long can that go?
What about a project that takes $100k in tokens and generates $200MM? The difficult part is determining how much was AI is the multiplier.
We are finding both to be true while quantifying requires new accounting.
[2 comments hidden]
However, that's very different from building and maintaining a product that will be used for live customers, where you have to consider security implications. This is what we're talking about here for Meta and Microsoft, and it's what I do for my actual coding job.
[hidden]
There is a level of trust on the model, the harnesses, and the operators that is required to move quickly.
To build that trust we are building systems and processes and getting the mindset that trust is the speed limit.
[hidden]
Then they tightened it down again and again. Right now, it's at $2k/month or $24k/year and the model restrictions are and to go wider.
Still, most employees are not hitting the limit but enough are that there's an exception process
[17 comments hidden]
[8 comments hidden]
In 2026 their revenue has gone up by a factor of more than 10x, and they no longer have just two whale customers.
I heard a rumor recently that customers spending less than $100m/year aren't even considered their "top tier" now.
[7 comments hidden]
Where are you getting the 2026 figures from? The Bloomberg article that was "on track to generate" and just prediction?
[6 comments hidden]
The most recent reporting from Reuters themselves (somehow not included in their more recent article about the IPO stuff): https://www.reuters.com/business/anthropic-ipo-valuation-hin...
> The company's financial trajectory already shows how quickly that equation is changing. Anthropic's revenue run rate was about $9 billion at the end of 2025, according to the company, before rising to more than $47 billion by May. Anthropic has projected revenue of at least $10.9 billion for the second quarter of 2026, more than double the previous quarter, on track for its first quarterly operating profit of $559 million.
Here's the FT: https://www.ft.com/content/4564e6a5-69e9-40a6-bf0f-a888f2f4f... - "Anthropic tells investors it will be profitable for second straight quarter"
> The company has told a small group of shareholders that its adjusted operating income will be positive for the second consecutive quarter, according to multiple people with knowledge of the matter.
These are leaked and self-reported numbers, but no matter how much skepticism you pile on them it still looks likely that Anthropic in 2026 have had some of the fastest revenue growth of any company in history.
[5 comments hidden]
If you're able to link something non-paywalled, I'd be happy to read thru.
[4 comments hidden]
They may be self-reported numbers, but they're consistent across multiple reporting sources. If Anthropic are lying to their investors about these numbers they will be in very real trouble with the SEC come IPO time.
And even if they're using funky accounting tricks to exaggerate their profits and downplay some of their losses, I expect the numbers for 2026 will still be really impressive.
[3 comments hidden]
Is the hope for them that they'll outpace their run with a massive revenue growth? I believe that they have that, but numbers like a net loss of $46b in 2025[0] seem hard to overcome.
Re: self-reporting, I'd be inclined to think they're using non-GAAP practices to make it sound more favorable, but I am by no means an expert, so willing to believe I am wrong.
I'll read the article you sent; this is just my off-the-cuff thinking.
[2 comments hidden]
I really don't think the 2025 figures are interesting at this point. Everything changed for them in 2026 - nobody was spending $1000/month/employee in 2025, there wasn't enough interesting token-heavy stuff to do with the models.
I think Anthropic might actually make it to profitability - they're earning more revenue than OpenAI and they've spent significantly less, too.
[hidden]
I do think their definition of profitability is likely skewed, but you're right that 2025 is out-of-date.
I did realize I forgot to link the article sourcing my claim in my above comment, and too late to edit, so here it is: https://www.reuters.com/business/finance/anthropics-ipo-pros...
[2 comments hidden]
However Elon was able to ipo before them which took a lot liquidity of the market. Their financials are exposed. The market condition and sentiment now is in the gutter. It would be very interesting to see how these would pan out
[22 comments hidden]
what. I can see a team of 20 costing 100k per month (but rare), but per person?
[20 comments hidden]
[9 comments hidden]
I suspect that SaaS is silently being eaten alive as an industry.
[2 comments hidden]
[5 comments hidden]
[3 comments hidden]
[hidden]
[hidden]
[hidden]
AI is very good at creating small codebases. It still sucks shit at maintaining huge ones where someone needs to understand how the codebase works.
I've seen what happens when companies just vibe code an entire mega codebase and it's really bad. AI is fucking terrible at grand architecture decisions.
[7 comments hidden]
I also doubt a longterm 2x multiplier for most developers.
[5 comments hidden]
[hidden]
[2 comments hidden]
[hidden]
[hidden]
It’s a contrived example, but to me it’s not improbable that a lot of the spend comes from tokens that are used very inefficiently. After all developers were pushed to try using LLMs without looking at costs.
I’m not saying the others are wrong, just that you can only tell something about marginal cost effectiveness, not total cost effectiveness of LLMs by measures to reduce AI spending
[hidden]
[9 comments hidden]
[5 comments hidden]
You can also read this as diminishing returns / AI isn't good enough, etc, but the simplest explanation is that they don't want to send money to Anthropic.
[hidden]
[hidden]
[5 comments hidden]
[2 comments hidden]
$1.4B/year is not a small number, even at Meta's scale, when it's money going to a competitor
[hidden]
[5 comments hidden]
1) Giving tight budgets for AI use after a few years of a total free for all, and
2) creating internal API marketplaces where one can access all the models including open weight and “Chinese” models.
This is increasingly sending tokens to the lowest bidder, which is a terrible setup for the big labs and their hopes of being worth trillions.
The big labs are tying to spin up “enterprise sales teams” like the traditional big SaaS players and trying to secure large contracts, but it’s broadly blowing up in their face and turning into these realtime token marketplace models.
[4 comments hidden]
[2 comments hidden]
[hidden]
A lot of developers spend significantly little time coding. So yeah, their coding time has decreased, but they were never coding for 6 hours a day every day to begin with.
[4 comments hidden]
[13 comments hidden]
[3 comments hidden]
Windows 11 shipped with a broken task bar, couldn't even ungroup items. No power users was ever involved in this.
[hidden]
[hidden]
Does anyone else remember this?
[hidden]
[hidden]
I use the Github Copilot app at work but it's really just a tool that gives models access to our Github repos and my local clones of them, and it's been super helpful to just tell it to reference other repos for patterns or integrations or whatever, get agents and an orchestrator to look at multiple repos to make a change across several projects and then implement them, etc.
And yeah you can use a selection of whatever model, I was using Opus 4.6 for a while, now I'm on GPT-6 Sol. It's been a very good experience.
[8 comments hidden]
[5 comments hidden]
Must be nice. At Sony we use Copilot (as an MS enterprise customer) and restrictions to token usage circa June have resulted in a hunger games of who-can-use-them-all-first, with the pool usually drying up by the 10th day of each month.
As for usage breakdown: we got a manually-published table most days just showing per-user token usage, until eventually we were told even that was somehow too difficult to provide.
Between that and learning how much the whole of Sony Corporate Group spends on Copilot (spoiler: it's hilariously little), I've sort of run out of disbelief.
[2 comments hidden]
[2 comments hidden]
I could see finance looking at it from a cost perspective if it was API costs, but sharing it with the rest of the users feels like a different measurement...
[hidden]
> I could see finance looking at it from a cost perspective if it was API costs, but sharing it with the rest of the users feels like a different measurement...
It is indeed surreal. I am not bullish about this company's technological capabilities going forward tbh.
[2 comments hidden]
Not sure where you see that. The API prices of Opus are cheaper than Astra and there was a study suggesting that the Anthropic abos are offering more bang for the buck:
[hidden]
Our accountants want that, our devs want that, but we cant have that. we would even be willing to pay a premium of 2x or 3x the price to have the enterprise benefits of private data that isnt retained. We would even be OK with all of the bottom of the barrel capacity and availability QoS too. We also would be ok Signing 3 year and 5 year lock in agreements.
We would be ok having a delay to the latest model by 30 days if that helps.
If anyone in Anthropic is reading. PLEASE give enterprises something like this. We want the Pro, Max 5x, max 20x plans. Charge us 3x the price for those exact plans but with the enterprise data privacy.
[3 comments hidden]
Is most of it spent on internal administrative tools? Is Microsoft cutting contracts with other SaaS vendors because their employees use AI to build their own alternatives? Is it just everyone querying the same questions to cut through the corporate drudgery and documentation hell? As an ignorant outsider, Windows is still terrible. Github is very mid. VS Code is... there. SQL Server/Windows server is ... still there
What about Meta? Instagram is the same as its been for years. WhatsApp the same. Meta RayBan app the same. Facebook.com the same. I don't use Snap, so is Anthropic's 2nd largest customer improving Snapchat? I doubt it.
That's my question to these large enterprises.
[hidden]
[hidden]
To also add an important context to the drop in Claude code users reported for meta - this elides that we are absolutely still using the Anthropic models full steam ahead, but are moving towards internal interfaces and harnesses.
So I would question if that 50% drop in CC users is more of an interface change than anything else.
I for one have stopped explicitly using codex or CC entirely. But the interfaces I am using still use those harnesses under the hood. I wonder how that is counted.
[hidden]
[hidden]
I wonder if these still exist, and if so, is MSFT worried about leaking them if they let their devs use other vendor's models.
[hidden]
[2 comments hidden]
And what were they doing with $100,000/month? I'd really like to know.
[hidden]
[hidden]
[5 comments hidden]
[hidden]
[2 comments hidden]
Or maybe not, but the bar was set pretty low that going all-in on AI might have been worth it.
[hidden]
[12 comments hidden]
[7 comments hidden]
There are people out there building AI building orchestrators for orchestrators for orchestrators for agents. The author of that blog post later claimed to be spending the equivalent of $122k/month on tokens (by rotating their usage between 21 accounts).
As far as I can tell, the only thing that this level of spend has produced so far is an indie 2D RPG video game.
[4 comments hidden]
Regarding Gas Town, see also https://yegge.ai/essays/the-shape-of-things-to-come/ :
> Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw
I hadn't seen the $122K figure mentioned previously. $87K for API-style pricing was mentioned in the above post, and ~$2,800/mo for multiple Claude Max accounts:
> My solution has been to create a token tap on $200 Max accounts, which for me work out to ~30x the list-price equivalent. So in reality I'm only spending about $2800/month out of pocket for my $87k "worth" of tokens. Though that number keeps growing alarmingly.
(He never seemed to provide numbers for Gas Town initially, whether what he paid, or what the API-style pricing would have charged, so it was interesting to get some actual figures. Sounds like a lot to me, but if he's genuinely getting multiple people's-worth of work out of it, then...)
(Also in the article: a little morsel of Emacs content, which was nice to see.)
[3 comments hidden]
[hidden]
[hidden]
Whether that produced anything worthwhile can be controversial, depending on the very subjective ways people quantify these things. All I'll say is that they have 18K and 27.7K stars on GitHub and apparently a comparable number of users.
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
The cybersecuritynews.com news one simply republishes details of a story published by The Information. At least they have the decency to LINK to that Information story:
https://www.theinformation.com/articles/meta-microsoft-work-...
... and of course the Information story is behind a paywall.
[hidden]
[hidden]
I like the tool. I don’t like how it’s the only tool all the time and has supplanted human communication.
[hidden]
[hidden]
[hidden]
[hidden]
I assume this is being pounced on by "I told you so" AI skeptics. Sorry but it's not what you were looking for.
[hidden]
[hidden]
[2 comments hidden]
What happened to last year's 'token maximising or you're fired'?
[4 comments hidden]
I am only half joking, I heard something like "my son or nephew did this cool thing with $X so we'll take $this_radical_step because of it" enough times over my career.
[2 comments hidden]
If you are used to talking to opus5.5 medium, going to GPT6.1 luna low will feel like a step down. Why would any employee take the (personal) risk?
[7 comments hidden]
[24 comments hidden]
[3 comments hidden]
[hidden]
[hidden]
[hidden]
After all, they should know how to compile their software. Any automation of that process is cheating their employer.
[hidden]
If my employer is putting scoreboards to see and champion who uses a tool which costs money to use, they shall pay for the tool.
Sorry, I'm not a ladder climber, yet I'm not mindless enough to bankrupt myself.
[2 comments hidden]
[hidden]
I can't fathom the logic of paying 10k a month for a developer and giving them an 800 device.
Insane. Your leadership is severely broken. If you can, work somewhere that values you.
[3 comments hidden]
[hidden]
Pretty early in my software career I had a job at a very awful company [0]. They wanted me to use the software "Weblog Analyzer" to analyze Apache logs and create reports that no one actually read. I was to do this twice a week. It wasn't hard and it only took like fifteen minutes so I didn't really mind doing it.
Initially I had a 30 or 60 day trial for the software, but once that expired the software (understandably) stopped working. I went to the director of technology, asking him to buy me to the software, and he responded back with "whoooa, hold up, that's not within the budget". To be clear, this software was, I believe, less than $100.
I responded back with something like "Ok, well then I'm not going to make these reports anymore". He acted surprised, like I was starting some kind of fucking mutiny, and he asked me why I'm being difficult. I said something to the effect of "I don't make that much money as it is, and I'm not going to spend my own money to make your company more money". [1]
He responded back with "Well [my direct manager] never asked me to buy this software", to which I responded with "Well [manager] is a moron who doesn't know anything".
Like a fucking sitcom, unbeknownst to me my manager was right behind me and heard the entire conversation. The next day I was fired. I don't completely blame them, it was admittedly very mean to say that about my manager even though I did (and still do) believe it.
I later found out that my manager had actually asked for the software years before, the director had said no and forgotten, and my direct manager had just bought the software with their own money.
Point is, at least some employers seem to have no issues making employees pay for their own tools, even for tools that they very specifically want the employees to use.
[0] This was a W2 job, not a contract gig.
[1] I even offered to try and find some alternative free software that might give us something similar, but he immediately poopoo'd that idea because they really liked these very specific charts that they never actually read.
[4 comments hidden]
[hidden]
An LLM is not too dissimilar to a Work Laptop or an IDE license.
[hidden]
I think that gets legally murky, if the employee is the one who pays for the tool.
ashleyn[78 comments hidden]
tinza123[8 comments hidden]
ihuman[7 comments hidden]
NewJazz[6 comments hidden]
therein[hidden]
ihuman[2 comments hidden]
98codes[hidden]
Zambyte[2 comments hidden]
wolvoleo[hidden]
But it was never a SOTA model.
wffurr[10 comments hidden]
nwhnwh[5 comments hidden]
CamperBob2[4 comments hidden]
nwhnwh[3 comments hidden]
netsharc[2 comments hidden]
nwhnwh[hidden]
coef2[hidden]
kelvinjps10[hidden]
actionfromafar[hidden]
casta[hidden]
ActorNightly[8 comments hidden]
Like its a no brainer to force your employees to use your own models, then RL train them to be better.
InsideOutSanta[hidden]
user43928[hidden]
I am not convinced that's the case.
p1necone[5 comments hidden]
locknitpicker[2 comments hidden]
Is security no longer a concern?
When you are using a third party model, you are literally feeding it not only your current software but also all the context and work fronts under development.
This is way more than granting a third party access to your internals. This is feeding it in advance updates on all their operations in real time.
NewLogic[hidden]
pimeys[2 comments hidden]
p1necone[hidden]
nateglims[29 comments hidden]
UltraSane[hidden]
seanmcdirmid[6 comments hidden]
claysmithr[5 comments hidden]
onion2k[3 comments hidden]
red-iron-pine[2 comments hidden]
ask it for recipes, weather, directions, explain baroque art, etc.
arguably still a common use case...
kingstnap[hidden]
Like in my experience the real burn is if there is some sort of feedback edge where an agent is producing stuff that eventually it has to reconsume.
Like if there is a dag of agents doing something its almost fine.
But if there is a loop somewhere without a huge amount of damping then you have an issue.
If you have ever been a human debugging in the middle of an agent it produces like so many commands and tests for you to do. Like way more than a human would.
I'm pretty sure they do the same thing to each other. Like the loop will amplify the yapping.
seanmcdirmid[hidden]
gradus_ad[hidden]
xpct[9 comments hidden]
iJohnDoe[8 comments hidden]
If it’s not already a thing, it will be soon, the mental health aspects of all professions moving at unsustainable speeds and what that will do to people.
pdimitar[4 comments hidden]
Nobody will take care of us. Ever.
ponector[3 comments hidden]
red-iron-pine[hidden]
pdimitar[hidden]
locknitpicker[2 comments hidden]
I don't think any vague claim of "speed" has anything to do with fatigue.
I do think the parallel world monitoring and validating N work fronts, followed by rounds of fixes where you lay there waiting for the agents to converge, is the defining factor.
Instead of you hunkering down and hammering out a single task where your full attention can be focused 100% on a problem, your mindset is a kin to juggling N balls and hoping to not let any of which to fall.
So it's not really about speed but throughput, and keeping up with the throughout rate is mentally taxing.
RataNova[hidden]
inSumErgoCogito[hidden]
devin[9 comments hidden]
SlightlyLeftPad[hidden]
Muromec[hidden]
mepiethree[6 comments hidden]
nradov[2 comments hidden]
jasomill[hidden]
Case in point: VSCode has a multistage "upgrade project to the latest .NET runtime with Copilot" process that takes several minutes of processing punctuated by frequent authorization prompts for even the most trivial project, resulting in the functional equivalent of
along the happy path.It has many more features than my shell command, of course, but none that I care about when migrating trivial projects from runtime version n-1 that are unlikely to hit breaking runtime or compiler changes that don't elicit warnings from the toolchain.
pjerem[hidden]
etempleton[2 comments hidden]
devin[hidden]
Jach[hidden]
joshstrange[hidden]
Yep, this is happening in lots of companies. I know my small company has about ~1/3rd of the developers working on one (different ones, all semi-personal projects).
I sometimes wonder if we are drawn to building these because it feels like one of the remaining challenges and a desire not to be an "LLM text shuttle". Additionally/alternatively, you can go as fast as you want with building a software factory and you aren't held back by the "bottlenecks" (PR review, QA, etc).
Working on my software factory is the closest to a "flow state" I've been able to achieve since LLMs turned the corner earlier this year and replaced the vast majority of our code-writing work.
mawadev[5 comments hidden]
Just one unsanitized input and you leak info. Or one hidden character and code may or may not belong to you anymore. Its very odd on many levels
user43928[4 comments hidden]
At best you produce some noisy signals that are going to have a tiny impact if even that.
And that's on a personal plan where you didn't opt out of sharing usage data.
Business plans offer zero data retention. This is a non-issue.
plasticchris[2 comments hidden]
user43928[hidden]
I think this data is likely worthless compared to curated RL tasks.
Semon132[hidden]
glerk[hidden]
trueno[9 comments hidden]
andybak[5 comments hidden]
Terr_[2 comments hidden]
andybak[hidden]
fragmede[2 comments hidden]
andybak[hidden]
compiler-guy[2 comments hidden]
... in 1988.
https://en.wikipedia.org/wiki/Eating_your_own_dog_food
vjvjvjvjghv[hidden]
latentsea[hidden]
BearOso[2 comments hidden]
iririririr[hidden]
wolvoleo[4 comments hidden]
And meta is only arguably so. Not really in the same league as OpenAI and Anthropic. More second tier like Google and xAi. (and of those Google is pretty close to the top two at times)
locknitpicker[3 comments hidden]
Microsoft's MAI-Code-1.1-Flash is on par with OpenAI's Luna line of models. If not for OpenAI's recent radical change of heart on Luna's pricing to hastily slap a 50% discount, MAI-Code-1.1-Flash could very well be the dominant cheap model.
insane_dreamer[2 comments hidden]
I see MSFT hasn't gotten any better at product naming :/
diegolas[hidden]
syngrog66[hidden]