HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

Claude Haiku 5.5

999 pointsby sfkgtbor 23 hours ago471 comments

Discussion

Loading discussion
  • TheAmazingRace · 23 hours ago

    I wonder if we have an AI LLM equivalent to Moore's Law. Like how often do we expect improvement in this technology and with what timing?

    • himata4113 · 23 hours ago

      double the information density every 2 days? serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.

      • qeternity · 23 hours ago

        Knowledge will be shifted to systems like n-gram augmentation which are relatively cheap and will not compete with reasoning capabilities for weight saturation.

    • dyauspitr · 23 hours ago

      Hopefully enough runway for an existing model to train the next to be better than itself with absolutely no human intervention.

    • bravetraveler · 23 hours ago

      I've heard tell about 100% of certain types of work being ended in batches of six months. For years. Truthfully, I'm skeptical, but accuracy wasn't prioritized.

    • ChaseRensberger · 23 hours ago

      reminds me of this blog post: https://campedersen.com/singularity

      • kator · 19 hours ago

        Whew, at least I won't have to hand-code solutions to the 2K38 problem!

    • onlyrealcuzzo · 22 hours ago

      Yes -> every 18 months they've gotten 90% more efficient for the same level of quality for about 5 years. There's little sign that trend is slowing. If anything, there's reason to believe that System 1 models (plus potentially 1-2-3 workflows) may increase that over the next 3-5 years. You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc. We haven't yet seen that at any size AFAIK.

      • thefourthchime · 22 hours ago

        Andrej Karpathy said once that he expects superintelligence could fit in 1 billion parameters.

        • onlyrealcuzzo · 22 hours ago

          Super intelligence that doesn't have to deal with the real world, maybe. I wouldn't be surprised if less than 1B param equivalent of our brain deals with solving math and writing computer programs and physics and all the things we tend to associate with "intelligence" - especially if you ultra optimized for that, I doubt our brain works like that. Dealing with the real world, I highly highly doubt it.

          • jstummbillig · 22 hours ago

            How about if we get away from written text as the input, to something more fundamental, that then also is able to produce text (among other things)? Given that humans learn to talk while having encountered a measly number of word instances, and, given enough time, we should always be able to improve on the lottery that is biology, it does seems fairly likely.

          • Gigachad · 19 hours ago

            It would be interesting if running ends up being a more complex task than advanced math. And our brains are just 95% allocated to dealing with the real world.

        • bakies · 16 hours ago

          What's super intelligence?

          • qznc · 11 hours ago

            Higher intelligence than any human. AGI (general) is about matching humans and ASI (super) is about surpassing humans.

            • bakies · 2 hours ago

              What useless definitions, ok. Not gonna take opinions of people using these terms too highly.

    • istjohn · 22 hours ago

      According to Epoch AI: > The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0] 0. https://epoch.ai/publications/the-plunging-price-of-thought

      • FooBarWidget · 22 hours ago

        Then why are AI plans still so super expensive, and AI spending going through the roof, while all the subsidies are ending?

        • jstummbillig · 22 hours ago

          Because it's increasingly useful and the thing you are substituting (human time) is much more expensive.

        • teaearlgraycold · 22 hours ago

          At least for me the Claude plans seem like an incredible deal and I never hit my limit.

        • adgjlsfhk1 · 22 hours ago

          The cost per fixed level of intelligence is dropping, but we're also getting dramatically more intelligent models.

          • verdverm · 17 hours ago

            MiMo-2.6 RL'd for ~$3.5M (not B), both main and flash combined, that is dramatically less and top 10 on https://artificialanalysis.ai/ https://mimo.xiaomi.com/mimo-v2-6 A frontier Ai is cheaper to make than a single 5/6th gen fighter jet, and maybe every fighter jet at this point.

            • JacobAsmuth · 16 hours ago

              Especially if you have millions of Opus 5.5 examples to train off of!

              • verdverm · 15 hours ago

                so tiring... you don't get to frontier from traces alone... also, who cares, the world is a better place if there are more awesome models at cheaper prices built with more efficient means we used to celebrate this kind of advancement, now it seems like astroturfing and belittling are the cool thing de jour

        • jrflo · 21 hours ago

          Because models are only getting better at a rate of 10% per year, people always want the best quality possible. You can get SotA performance from a year ago for a fraction of the cost, but why would you use Opus 4.5 when you can use Opus 5.5?

        • srdjanr · 21 hours ago

          Apart from what others said about using more intelligent models instead of cheaper ones, token usage is also increasing a lot. Classic Jevons paradox

        • stephbook · 20 hours ago

          https://en.wikipedia.org/wiki/Jevons_paradox AI gets cheaper, people use it everywhere. Google searches, for example. Now we want to crack math problems and spend weeks with unreleased models. If you used GPT-2, it'd be incredibly cheap. You basically can't use it for anything and it's simple to serve.

          • versteegen · 17 hours ago

            You mean if you used a modern LLM of GPT-2-level quality. Vanilla transformers like GPT 2 are ridiculously inefficient in comparison.

        • f6v · 19 hours ago

          Reddit is full of people complaining how they burn their 200$ sub in half an hour by starting ten Max sub agents. That’s to say, many people just don’t know what they’re doing.

        • BenzeneDream · 16 hours ago

          To pay for the training of the models which are getting bigger and more expensive. So intelligence is getting cheaper overall but the need for ever-increasing intelligence can't be sated.'

  • sfkgtbor · 23 hours ago

    Happy about the Sonnet cache read price cut.

    • minimaxir · 23 hours ago

      That was effectively required to match GPT-6.1 Sol (costs and caching prices are now equal). Sonnet 5.5 made zero sense to use over Opus 5.5 under the old cache prices.

  • minimaxir · 23 hours ago

    Pricing is...a bit weird. Input $0.10 / MTok for prompts up to 100,000 tokens $0.50 / MTok for prompts over 100,000 tokens Output $0.50 / MTok for prompts up to 100,000 tokens $2.50 / MTok for prompts over 100,000 tokens 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use. In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])

    • giancarlostoro · 23 hours ago

      I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x). I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.

    • tr4656 · 23 hours ago

      Luna does as well, but just at a higher limit. From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

      • minimaxir · 22 hours ago

        Huh, that disclaimer is on the model page ( https://developers.openai.com/api/docs/models/gpt-6-luna ) but not the pricing page. Annoying. Fixed.

      • tripleee · 22 hours ago

        So even at the 1.5x/2x rate luna is still half the price of this. Weird pricing strategy from Anthropic. I'm sticking with Luna if I don't need a super smart model

        • usef- · 20 hours ago

          You're judging purely by token cost I assume, not cost per completed task? The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.

          • tripleee · 20 hours ago

            yes, that's true. I should be looking at the $/completed task

          • RussianCow · 19 hours ago

            The cost per task from Artificial Analysis is roughly 3x higher at every reasoning effort level for Haiku than Luna. Sol 6.1 on medium has the same cost per task as Haiku with significantly higher intelligence. According to those numbers (which you should take with a grain of salt), from a pure cost vs intelligence standpoint, you're better off using Luna for economics and Sol for intelligence. With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.) https://artificialanalysis.ai/models/releases/comparisons/cl...

    • j45 · 23 hours ago

      It could be to incentivize people to not be lazy users of tokens.

    • Eridrus · 23 hours ago

      It's actually existing flat per-token pricing that is weird. Neither encode nor decode are linear in compute, so providers need to price for average expected length. This is just getting closer to the true cost of generating tokens.

      • sebzim4500 · 21 hours ago

        Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO

        • stingraycharles · 14 hours ago

          I think this was originally started by Google who first offered these large context windows and others just followed suit.

      • foota · 21 hours ago

        My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.

      • hgoel · 20 hours ago

        Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.

        • vardalab · 14 hours ago

          You would think with all this AI available, they could figure out the pricing to be as granular as necessary.

          • anthonypasq96 · 11 hours ago

            people dont like when the price of something is a giant formula or ???

      • willsmith72 · 8 hours ago

        they don't need to, companies smooth out costs all the time. cost of goods sold isn't a great way to price software products

        • stingraycharles · 7 hours ago

          Yeah, also, the real cost here is memory, not compute. Which is why KV cache quantization is a thing.

    • insanitybit · 22 hours ago

      I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.

    • dannyw · 22 hours ago

      Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here. For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc. These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k. In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.

      • WinstonSmith84 · 22 hours ago

        noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ...

        • goosejuice · 17 hours ago

          No? Don't use these lower end models to work on small well defined tasks with a frontier model orchestrating? This approach works very well for me and don't have any issue staying under 100k. I have no idea if it's cheaper but it does seem to be much faster for tasks like QA.

      • JacobAsmuth · 20 hours ago

        The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA.

      • flockonus · 20 hours ago

        > it’s really impressive how much intelligence per dollar has grown in just a few short months. Open weights models giving a distant salute from afar

        • pimeys · 19 hours ago

          Yes, it was weird to see MiMo and DeepSeek missing in the article's comparison...

          • RussianCow · 19 hours ago

            It's not that weird. Most companies considering paying Anthropic are probably not considering Chinese models as alternatives. Many don't even realize they exist.

            • doodlesdev · 18 hours ago

              The thing is: availability of near-SOTA cheap Chinese models is forcing OAI and Anthropic to bring prices down and offer more efficient models, instead of simply focusing on super expensive SOTA LLMs.

              • RussianCow · 17 hours ago

                Is it? I would guess that it's much more about the race to get customers as they both near IPO than anything to do with the Chinese models.

                • verdverm · 17 hours ago

                  we are actively preparing to move our devs from closed to open models, take it as a piece of anecdata the trend in industry is clear by now though

                  • pimeys · 12 hours ago

                    Yes. We did the same a few weeks ago. Everybody I talk with in the EU startup ecosystem is either doing the same or considering doing it.

            • epolanski · 18 hours ago

              "companies" is a meaningless metric. If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6.

              • tranceylc · 18 hours ago

                Vibe coding gta 6 haha

              • RussianCow · 17 hours ago

                I said "most companies considering paying Anthropic", which is not the same as "most companies". I also don't agree that "nobody" is doing this; I have lots of anecdata suggesting otherwise. Maybe the majority of companies are using the easy option of Copilot or Gemini like you said, but it's nowhere near 99%.

                • lifeisloving · 15 hours ago

                  Yes it is. Literally every company Ive spoken too about this only has MS teams co-pilot lol. Like all of them are using the most basic form of AI. They dont even know about Anthropic and think its just "AI". By the way this includes one of the largest power systems design firms in the world, who helps build many datacenters... Uses only MS copilot.

                  • blackqueeriroh · 15 hours ago

                    Do you understand how many companies there are? Your anecdata is wholly useless

                    • vardalab · 14 hours ago

                      Well, I can add two more points to your anecdata. I personally know of two companies where only exposure for peasants is Copilot.

                      • drob518 · 6 hours ago

                        Yep, this doesn’t surprise me.

                    • lifeisloving · 14 hours ago

                      Umm, just go look up the whole of software revenues globally and compare that to OpenAI. Its very small. Nobody is paying for SOTA models outside software engineering and ancillaries, except for some outliers fields like consulting. Then there's techies in every field who use it, but not companies themselves. Companies as a whole, do not care about this technology, except tech companies and its adoption cant even be compared to CRMs. A company might buy a SaaS product with AI but most of them are not purchasing Anthropic subscriptions lol.

          • stavros · 18 hours ago

            Is Mimo good? I've never tried it, but I've seen it mentioned three times in this subthread alone. DSv4.1 is my daily driver.

            • vardalab · 14 hours ago

              As context grows, it gets slower and dramatically slower. Otherwise it's pretty good.

            • sausagefeet · 9 hours ago

              I have been settling on using Mimo 2.6 flash for discovery/research and glm 5.3 flash for implementation. So far I've been pretty happy with the results, but mimo can be quite slow. If speed is an issue, I just switch to glm 5.3 flash.

          • drob518 · 7 hours ago

            And GLM 5.3 Flash.

            • pimeys · 4 hours ago

              I haven't really tried it that much. I use GLM 5.3 a lot, especially with Coralbricks where the cache reads are free so I can let it think a lot without breaking the bank in the following turns. It's really good for bughunt, planning, and security work. Maybe I should try the flash...

      • RussianCow · 19 hours ago

        I said this in another comment, but Artificial Analysis has the cost per task of Haiku on max roughly equal to that of Sol on medium, and the latter is significantly more intelligent. (And I'd wager that Sol probably finishes tasks more quickly, even with Haiku inference being faster.) So Haiku really only makes sense on lower reasoning levels, and only if you care about intelligence and speed more than you do about cost effectiveness (where Luna currently dominates). And that's without even bringing Chinese models into the mix.

    • system2 · 22 hours ago

      Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?

      • user43928 · 22 hours ago

        Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users.

      • wyrdcurt · 22 hours ago

        Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash for almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money.

        • pimeys · 19 hours ago

          Luna is not really the fastest. You need to use it in high/max to get the good output for what it is good for: summarizing. And that is already close to two minutes per task...

        • girvo · 19 hours ago

          I’m on the Legacy v2 plan and same: nothing comes close to 5.3 Flash’s value on it. It’s crazy, no wonder they discontinued them!

        • RideOnTime22 · 15 hours ago

          There's also fomo and what I believe is faniticism. Even for simple tasks why use X if I know "Y Max" is available and on paper, better? And why use something else when your favorite company releases something. Surely it must always be the best one to use.

        • vardalab · 14 hours ago

          z.ai speed has been horrific, worse than sol-6.1 was last week, I am not renewing my sub once it expires.

      • mrngld · 22 hours ago

        That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there.

        • RussianCow · 19 hours ago

          Except for the new MiMo V2.6 models, which appear to give some of the best value right now, at least on paper. (I haven't tried them so I can't speak from experience.)

      • aesthesia · 22 hours ago

        There really aren't any models at 10% of the price of Luna or Haiku.

      • skeledrew · 22 hours ago

        Well, unless you're using OpenCode Go, it's per-token costs (even if already super low), while Haiku falls under the Claude sub. It's just more straight forward and you aren't feeling a "loss" with the sub.

      • pkulak · 20 hours ago

        Where do you get this 10% number? Checking providers I know/respect, and GLM 5.3 flash is $0.15/m. Haiku is $0.10/m.

      • usef- · 20 hours ago

        On subscription pricing a $20 Anthropic subscription gives >$500 equivalent tokens, which is not so different, and you get smarter models. API pricing has decent margins. And Opus 5.5 is really good.

      • ray_kay777 · 20 hours ago

        People who are stuck using Bedrock in-geo due to their company policy (me).

      • nharada · 19 hours ago

        Isn't the point of this release that it's comparable? AAI Index // Input // Output Haiku 5.5: 43 // $0.10 // $0.50 Mimo 2.6 Pro: 46 // $0.43 // $0.87 Mimo 2.6 Flash: 38 // $0.10 // $0.28 Seems competitive to me? Plus then I don't have to manage multiple providers

      • flaburgan · 11 hours ago

        They say it in the announcement: Haiku is basically useful to be called as a subagent. So you use Opus, and you want to investigate your production logs, instead of throwing them right away which is going to use a massive amount of token for mostly noise data, Opus asks Haiku to determine patterns to extract only the relevant logs and feed it back in Opus. The models are meant to be used together.

        • walthamstow · 6 hours ago

          > The models are meant to be used together. If that was the case, then Claude Code would use smaller models for subagents. It doesn't. The subagent always inherits the parent model unless you tell it specifically not to.

    • esafak · 22 hours ago

      It's their creative way of 'matching' Luna's prices.

    • AustinDev · 22 hours ago

      encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens. There are plenty of workflows like translations where you'd easily be under the cap.

    • enraged_camel · 22 hours ago

      >> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents Your vibes don't appear to be supported by facts. From the announcement: >> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.

      • Philpax · 22 hours ago

        People weren't using Haiku 4.5 for agents before. 5.5 is good enough that it might be.

      • StilesCrisis · 22 hours ago

        Haiku 4.5 users were using it for Kleenex requests because that was the best it could do.

        • enraged_camel · 21 hours ago

          Not really. We use Haiku 4.5 to turn users' natural language queries and requests into fairly complex structured specs for interior design and construction. It has near perfect accuracy.

          • dotancohen · 20 hours ago

            How many examples are in your prompt? How large is that prompt? Or do you have some other way of tuning the output? I'm asking to learn for a similar project, not to discount anything you're saying.

    • Tiberium · 22 hours ago

      There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one. You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc

      • AtNightWeCode · 22 hours ago

        > ...this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5. So, it is might be even worse.

        • Tiberium · 22 hours ago

          No, it's just Haiku 4.5 is so old that it predates the new Claude tokenizer change in Claude 4.7+

    • HarHarVeryFunny · 22 hours ago

      Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor. For this application 100K token input is plenty. Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.

      • martianvoid · 22 hours ago

        I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification

        • HarHarVeryFunny · 22 hours ago

          The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks. For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost. I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.

    • port3000 · 22 hours ago

      They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.

      • cogman10 · 20 hours ago

        I think they are also trying to make sure Deepseek and other chinese models don't eat their lunch. They need something price competitive.

    • mnicky · 22 hours ago

      You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget.

    • alexchamberlain · 21 hours ago

      Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low".

    • sixtyj · 20 hours ago

      Chatbot could be < 100k tokens.

    • mkotlikov · 20 hours ago

      If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models.

    • jeremyjh · 19 hours ago

      If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful. I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.

      • serf · 18 hours ago

        >If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. "less context is better and if you can't get stuff done with less yur bad" is the worst argument ever . it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images. it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.

        • jeremyjh · 17 hours ago

          It isn't an argument, and I never said it is better. It is an explanation for why the pricing break is relevant. Anyone can do the same things to take advantage of that pricing. Would you prefer I not explain a basic fact to someone who may not know what is possible? The services are priced this way because larger context has significantly higher costs. That is a fact about the technology and it is true for every provider. So moaning about it isn't useful. On the other hand, there are a lot of people who don't manage context effectively - who start every session with 60K tokens - and that is significantly hurting the performance of every single thing they do with coding agents.

          • komali2 · 15 hours ago

            So what's your setup? Speaking as a "open terminal in repo, open Claude, say 'do this thing please '" kinda guy I'm interested in learning about these more advanced techniques.

            • jeremyjh · 5 hours ago

              I use oh-my-pi (omp.sh) - mostly with stock settings and skills but its a "fully loaded" harness that you don't really have to add anything to. I do change a couple of things: I set it to prefer subagents, and I enable rewind. It is crazy that rewind is not a default, it saves a TON of context - if the agent goes down a crazy path that burns a lot of tokens it can rewind to an earlier checkpoint with the exploration or bug hunting summary. Presently I'm using sol-high for the default agent which does orchestration and a lot of smaller investigation and coding tasks itself. sol-max for planning and review. Luna-max for planned coding and general tasks. I also have a $10 minimax plan and use M3 for exploration and library roles, but I could probably be using Luna for that just as well and still only very rarely run into usage issues. I don't use any plugins or skill libraries apart from Caveman and I'm not sure how useful that really is anymore so I'd start without it so you have a baseline to compare. I do think it reduces context usage a bit but I haven't measured it recently. Caveman also includes some team, agent & investigation skills - again they might be helping but I haven't re-evaluated since like 90 days ago.

    • solenoid0937 · 18 hours ago

      This is pretty good tbh

    • judetechdevs · 15 hours ago

      interesting! thx for your info

    • chaostheory · 14 hours ago

      Going on a slight tangent, I've found that Anthropic (for my work) costs about 4x as much as OpenAI give or take, specifically Opus vs Astra

    • dj_io · 13 hours ago

      Haiku doesn't seem most cost effective solution. Jev like classifier can work in a fairly smaller model which are 1/10 of the cost. Opus is SOTA so I get the use case for one being restricted to Anthropic ecosystem. I think the use case for Haiku is mainly for users using on their chat for pro subscribers and free users to maximize their quota.

    • krzyk · 1 hour ago

      They want to get on the Luna market. If there was no Luna you would see only the >100k pricing, but because we have Luna, they had to lower price for something.

  • maz1b · 23 hours ago

    Wow, the rate of improvements in the AI era is staggering. GDPval-AA v2.1 as of now: 1620 GDPval-AA v2.1 for Haiku 4.5: 735 The 100k tokens pricing makes sense, looks to be a hedge against OpenAI's decisions API and Jev or its open source alternatives that are springing up. Nice release, congrats to Anthropic.

  • sroussey · 23 hours ago

    It’s about time Haiku got an update!

  • iagocc · 23 hours ago

    Where is Pelican? ehehhe

    • rvz · 23 hours ago

      [flagged]

      • InsideOutSanta · 22 hours ago

        So do you have the pelican or no?

        • rvz · 22 hours ago

          [flagged]

          • tomhow · 21 hours ago

            Can you please not be so sneery/grouchy? That's far worse for HN than suboptimal benchmarks. The guidelines specifically ask us to avoid being curmudgeonly.

          • InsideOutSanta · 20 hours ago

            Ok, so I looked at all of your links, but nary a pelican to be found. How am I supposed to know what all these numbers mean if there is no pelican?

          • Mashimo · 18 hours ago

            What is i want to code svg files though?

      • swalsh · 22 hours ago

        I think it's a joke at this point, but also the visual benchmark is a remarkably dense method for demonstrating how good a model is.

        • InsideOutSanta · 22 hours ago

          Yeah, people like to poop on the pelican. But pelican quality still correlated with overall model capabilities reasonably well, and you can immediately see and interpret it. It's a running gag, but it also does have some actual value.

  • TomGarden · 23 hours ago

    From these selected benchmarks, it looks like it smokes Luna capability-wise. Excited to put it through its paces

  • margorczynski · 23 hours ago

    How does the price compare to Luna? At least looking at the numbers it is noticeably better at most tasks.

    • onlyrealcuzzo · 23 hours ago

      IMO, this is better. Luna is super cheap, but it's not that capable. At higher levels of reasoning, it's not that fast. This is more expensive, but it also looks like it's better enough that it's far more useful. I also won't be surprised if you look at cost per completed task + wall clock time that it comes out ahead for the majority of what you'd want to actually use it for. Luna will still be a great option for doing non-engineering tasks super cheaply.

    • TomGarden · 23 hours ago

      For prompts under 100k tokens, it's priced the same as Luna - $0.10 in, $0.50 out. For prompts over 100k tokens it's 5 times more expensive - $0.50 in, $2.50 out.

  • simianwords · 23 hours ago

    I remember a friend asking me why LLMs suck so bad. She was using Haiku 4.5 and that poor model couldn't keep track of the context within 3 messages. She said she was using Haiku 4.5 because she was advised to be careful with the spending. I hate that model so much lol.

  • seaal · 23 hours ago

    The monthly API credits for Max plan seems fantastic, especially considering Haiku pricing. Being able to actually use my Claude plan for other harnesses and use-cases on top of regular CC usage is everything I wanted. Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.

    • 0gs · 22 hours ago

      yeah totally agree. esp how efficient it can be to have a subscription quota-paid orch spin up a bunch of API agents, this is kind of like free money to encourage what was already an easy way to save money (via batch pricing)

    • thepasch · 22 hours ago

      Note that this is Anthropic Trojan-Horsing the previously announced June change in with a model release, where the Claude Agent SDK can no longer be used with Claude subscriptions and is now billed with API credits only. https://support.claude.com/en/articles/15036540-use-the-clau...

      • InsideOutSanta · 22 hours ago

        Ah, that sucks. I'm using Paseo to run Claude Code; I guess that just got a whole lot more complicated.

        • cjav_dev · 18 hours ago

          We just updated the docs to clarify how API credits can be used w/ claude -p: https://support.claude.com/en/articles/15036540-use-the-clau...

          • InsideOutSanta · 8 hours ago

            Thank you, that's great news! "You can still use the Claude Agent SDK, claude -p, and third-party apps with your subscription limits."

        • antoniojtorres · 17 hours ago

          Using Paseo as well. This would be a deal breaker for me for sure.

      • dr_kiszonka · 19 hours ago

        Thanks for sharing this! These vendor lock-in attempts are very annoying.

      • nr378 · 19 hours ago

        Yep they're definitely getting ready to yank using your subscription with the agent SDK - this page has just been pulled: https://support.claude.com/en/articles/15036540-use-the-clau... That's not a good sign for Conductor...

        • cjav_dev · 18 hours ago

          this is a common question. we're updating the faq now to make sure it's more clear Update: We just updated the docs to clarify how API credits can be used w/ claude -p: https://support.claude.com/en/articles/15036540-use-the-clau...

      • luketaylor · 18 hours ago

        Sorry for the misleading wording on this page; we’ve updated it to clarify that the Claude Agent SDK can still be used with subscriptions.

    • skeledrew · 21 hours ago

      > Being able to actually use my Claude plan for other harnesses Wait what? This has gotten their blessing?

      • neucoas · 20 hours ago

        You could always use Claude models on other harnesses via API... just not via subscription. Now they give you $100 worth of API tokens to use on opencode or Pi. Which is better, but still not the same as OpenAI were you can use the subscription on Pi without problems.

        • dfdydx · 9 hours ago

          So you have the normal subscription usage plus an extra $100 in API tokens now?

      • copperx · 20 hours ago

        Absolutely not.

    • laurels-marts · 18 hours ago

      OpenAI has been a disaster lately.

  • patrickwdaly · 23 hours ago

    How are y'all using Haiku though? I rarely select it.

    • svachalek · 22 hours ago

      Opus often picks it when it's doing a "find me something" subagent. But largely it's been held back by being fully a year old at this point, and priced at a much higher price than models that are far more capable.

    • steve_adams_86 · 22 hours ago

      It's great at parsing documents inexpensively. For the few skills/plugins I've made, I usually instruct Claude to use Haiku for low-reasoning grunt work.

    • mariocesar · 22 hours ago

      I have a zsh functions that calls claude code with haiku to suggest commit messages, is faster and the instructions are two lines. I also have an "ask" script that I use daily to ask simple stuff, it can access websearch and webfetch, it's more than enough to parse logs, ask for commands, quick research on the internet, small stuff. https://github.com/mariocesar/dotfiles/blob/main/common/.loc... I use haiku for things that needs to be quick, have really clear instructions.

      • tetraodonpuffer · 20 hours ago

        with claude -p seemingly now using api credits I guess this approach will have to change unfortunately :/ I wonder what will be the best cmdline way to do things like these

    • gghootch · 22 hours ago

      I was waiting for this. Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them ( https://evalship.com )

    • swalsh · 22 hours ago

      I've been using GPT-6 Luna in some capacity for nearly all my agent workflows. It's just a really good model, and the pricing is cheap. If Haiku 5.5 is better, and the same price (under 100k context... which is a big caveat) i'd probably swap it.

      • dannyw · 22 hours ago

        It’s absolutely better than Luna. It feels closer to a “sonnet 5.2” if that makes sense. Of course it’s not as big, and hence falls-off quicker. I’d consider the 100k a “promotional price” to match Luna’s token pricing while delivering noticeably more intelligence.

    • Plutoberth · 22 hours ago

      I'm building a game that incorporates LLMs as a game mechanic. I've been using Luna, but I'll probably switch to Haiku.

    • notatoad · 22 hours ago

      not haiku, but luna - last week i used it for things like "read this historical dump of 15k support tickets and break them into categories that make sense, then propose help docs that i could write to handle the most frequent queries in each category" used <10% of my 5hr limit on a $100 codex plan.

    • hector_vasquez · 22 hours ago

      My software application uses Haiku in production more or less as a Jev. I do not use it for coding or development.

      • mrkn1 · 20 hours ago

        Why not have a CPU-first decision model for free? check out gutsy

    • apothegm · 16 hours ago

      Translating GPT’s word salad to English at the end of an agent interaction.

  • tpoacher · 23 hours ago

    Good to see Anthropic back alternative OSes.

    • Ectiseethe · 10 hours ago

      Poems, then BeOS, then a model named for both. Same name, three lives.

  • afrnswrth · 23 hours ago

    The important question though...how does it do making a pelican on a bicycle?

  • caaqil · 22 hours ago

    > Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers. If you block pentest or "other techniques more likely to be used by attackers", then what does "permit a wider range of defensive tasks" even mean? Any defensive task that's meaningful is almost indistinguishable from legitimate red-teaming that then falls under 'likely to be used by attackers". If only they would just stop nerfing these models, that'd be great. No APT is waiting around for Anthropic's permission, so might as well let us have some cool stuff.

    • TuxSH · 22 hours ago

      Yeah it's too little too late, cat's out of the bag as people know that GLM 5.3 exists and is great at defensive and offensive cybersec. (sadly Mistral Large 4 isn't up to par - but Mistral serves GLM at 130 tps!)

  • simianwords · 22 hours ago

    > Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information, see our Help Center article. Did anyone read this? We get free API credits on some plans now

  • swalsh · 22 hours ago

    Top of the page in 17 minutes? Now I know what y'all do while your agents are working.

    • skeledrew · 21 hours ago

      It's a brave new world... of idleness!

    • niceguy1827 · 18 hours ago

      https://xkcd.com/303/

  • AtNightWeCode · 22 hours ago

    Probably the same scam as the last Haiku update I guess. Uses more tokens to compensate for the lower price.

    • AtNightWeCode · 21 hours ago

      To correct myself. The price was not lower. It was up about 20% for the tokens. But, the big price hike was that it used a lot more tokens for the same tasks.

  • charlesabarnes · 22 hours ago

    > Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes

    • geek_at · 22 hours ago

      This is literally for you to get tangled in their api and when they stop giving you the allowance they hope you will just continue to pay

      • enraged_camel · 22 hours ago

        What does "tangled in their api" mean? Switching is pretty easy.

        • charcircuit · 22 hours ago

          Not really, you have to fiddle with generating api keys and setting environment variables. Meanwhile with Anthropic it will just start charging you API prices for the tokens you are generating without even a single warning.

          • enraged_camel · 22 hours ago

            >> Not really, you have to fiddle with generating api keys and setting environment variables. That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.

            • charcircuit · 22 hours ago

              The user could have always done that regardless of if the user has the option to be charged API rates on or off.

          • ClikeX · 7 hours ago

            Of all the things I've had to deal with with migrating solutions and frameworks. Generating API keys and setting environment variables are the least time consuming of all of them.

      • tomjen3 · 22 hours ago

        That's an old tactic for an old world. You only need, what, half an hour with your agent of choice to write you out of that?

      • losvedir · 22 hours ago

        Nah, it's pretty trivial to switch providers (especially with Claude's help, ha). This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.

        • eli · 20 hours ago

          Or to discourage people from using cheap subscription tokens as part of automated workflows

    • tech234a · 22 hours ago

      OpenAI will probably add this to their plans within a week

      • alasano · 22 hours ago

        With OpenAI you can just use Oauth and get a token to use your subscription. Anthropic isn't even close to being this useful.

        • Iolaum · 22 hours ago

          Biggest reason for an OAI subscription instead of Ant imo. Biggest loss is that Ant models look like they are genuinely better.

          • matsz · 22 hours ago

            > Biggest loss is that Ant models look like they are genuinely better. This changes on a weekly basis, I ended up with subscriptions to most of the providers (except for X.ai).

            • floydnoel · 5 hours ago

              I even got one to X.ai because I wanted to compare them all. The models were fine but the usage quota on X was extremely low compared to OAI/Ant

    • thepasch · 22 hours ago

      This is them sneaking in taking the Claude Agent SDK (claude -p) off of subscription plans through the back door along with a model release. They previously wanted to do this in June, but backpedaled after huge backlash: https://support.claude.com/en/articles/15036540-use-the-clau...

      • sambaumann · 22 hours ago

        Even after the June changes there was some allowance to use agent SDK on the pro plan. This will move me to codex tomorrow if agent SDK is really blocked on pro

      • sanex · 22 hours ago

        Those mfers. I'm using this for work! I use my work teams plan with pi so I can do all kinds of custom workflows that I can't in Claude Code. Time to convince management I need OpenAI instead.

        • stsch · 21 hours ago

          Time to convince management (and yourself) to build some skills. :)

          • sanex · 21 hours ago

            Spent many years building skills, I'm just working on a different level now.

          • stavros · 17 hours ago

            Yeah there's no way in hell you're going to convince management to pay 10x for the same work because "human skills".

          • persedes · 16 hours ago

            Or just for the token usage. The rug pull is coming eventually

      • martinald · 22 hours ago

        Do we know if claude -p is now drawing from this API usage?