HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

GPT‑6 and Intelligent UI for everyone

729 pointsby joshuawright11 23 hours ago428 comments

Discussion

Loading discussion
  • throwaway7783 · 23 hours ago

    Is this OpenAI catching up with Anthropic artifacts? At the same time they say "We’ve trained GPT‑6 to compose responses using text, visuals and interactive elements...", rather than a harness.

    • davvie · 8 hours ago

      Looks similar but inline

  • solarkraft · 23 hours ago

    AFAIK, they already had a simpler form of this. It was kind of an obvious next step. Now connect it up to tools so that we can again comfortably do the things that are more precise by hand! The “pick a color” or “select the width on a slider” use case is coming closer.

  • Tiberium · 22 hours ago

    I find it extremely strange that they're adding GPT-6 Sol to Chat over GPT-6.1 Sol which is significantly more capable.

    • wincy · 22 hours ago

      I’ve hit model not available limits for GPT 6.1 Sol multiple times over the last week. It’s never for very long, and I can switch to Astra, but it seems like OpenAI is struggling with capacity. This has happened with my personal $200 Pro connection and my Codex enterprise connection.

      • 0xFluegel · 20 hours ago

        I had a codex session stop in the middle because the auto-reviewer timed out with something like "auto-review not available at the moment". The capacity issues are very clearly observable since gpt6-astra launched.

    • xpct · 22 hours ago

      I really don't like having to open Work sessions for one-off questions, only because they're limiting what models they put in the chat. The older models are just too dumb for some things.

      • zamadatix · 22 hours ago

        The whole Work/Chat/Codex split is maddening in the way it's implemented. It's a pain to switch back to the right project, it's a pain to switch forward, it's a pain to try to remember which chat was in what.

    • exitb · 22 hours ago

      Given the extreme short time between 6 Sol and 6.1 Sol, I suspect they don’t actually have much in common and 6.1 is a heavier model rebranded as Sol in a panic response to poor agentic capabilities of 6.

      • ismael_rr · 21 hours ago

        I suspect that 6 sol was a better version of 5.6 terra (and note that in the 6 sol and luna release, they took out terra, and price 6 sol at 5.6 terra pricing), then the backlash from lesser capabilities made them roll out 6.1 sol as the actual 5.6 sol - size modee.

    • MikhailTal · 22 hours ago

      6.1 is Astra minor. Way more capable but also way heavier+slower. Its really 2 different models, they just shipped it as sol to recover from the gpt6 disaster lunch, where they tried to pass terra 6(or a cheaper model) as sol but it was worse than expected

      • sscaryterry · 22 hours ago

        Anecdotally this is what I experienced, do you have sources for this?

        • spijdar · 21 hours ago

          It's purely circumstantial, but 6 Sol supports no reasoning (same as 6 Luna and 5.6 Sol/Luna), while both Astra and 6.1 Sol do not support "reasoning = none".

    • laurels-marts · 18 hours ago

      gpt-6.1-sol was released 7 days after gpt-6-sol only because they were bleeding to fable-5.1/opus-5.5 from Anthropic. It was a desperation move and certainly ate into their margins massively. Since their margins on gpt-6.1-sol are much slimmer than on gpt-6-sol they have to use it somehow to preserve compute and make profit. Frankly I think OpenAI is a mess and is playing catch-up with Anthropic… I don’t know how they will recover.

  • fang2hou · 22 hours ago

    I'm also building something similar for an internal project, based on the vercel's json-render design. It hasn't been too challenging, especially since the release of faster models like 5.6 Luna.

  • tamimio · 22 hours ago

    This looks bad, I want the output to be as much as text based so I can easily export it and further processing it, plus, this might make the resource-eating app even worse.

    • ardaakman · 22 hours ago

      Visuals are better for consumers/general public, when the use case is "help me cook X meal", or "where can I stop by to buy gas on the way to San Jose".

      • sshussain270 · 21 hours ago

        Rading most comments, people don't seem to be very impressed by it, I ain't either tbh - its not something mind-blowing (I have been doing this with Astra myself just by writing better prompts to make explainer interactive interfaces), but I do think its a good first step towards a better way of consuming information/answers than just reading text-vomits. I wonder if there is a way to train an LLM native to this kind of thing?

        • tripleee · 20 hours ago

          where are you reading the comments? Wasn't it just released? I can't imagine anyone I'd consider "general public users" to know about it yet

    • oh_no · 21 hours ago

      then just ask for text-only?

    • auyez · 12 hours ago

      I think this is really a future, this is the product that I really expected to become reality at some point. If we assume agents become even more common in the future, then I don't think we would really have a lot of modern software. Most of it is really is not needed at all. Most apps that people use are really very similar and don't have anything original. It would make more sense for it to be more personalized. For example someone might prefer not to interact with interface at all, and just access services just by chatting with a bot, while other person might want only last steps to be provided as UI. For example to see summary of his cart before paying. While another person might be insterested in just browsing all options. Some might prefer to have filters, while other would want AI to filter everything for them. I think the real result would be when all services would be automated by AI. Like if you want to be a small business, you no longer need to build anything. You just describe what real life services can you provide: delivery, barber, baker, cleaning, repair. Then all of that would be in some agent network, and available to other agents as an option. While humans on both sides would just get a bridge between them in whatever form is most comfortable for them. Buyer might get a list of bakeries as normal ecommerce website, while the baker might just be someone who receives phone calls explaining him what his next order is in human voice.

  • haute_cuisine · 22 hours ago

    In the video, they showed examples of ChatGPT making interactive tutorials on how to fold origami, how to arrange colours/interior and how to assemble a bike. Supposedly, people were struggling to follow written manuals and they needed an interactive explanations. I'm not a mathematician and I would certainly love having a tool that would do ELI5 on some complex stuff, but I'm really worrying about using this too often and outsourcing my ability to do stuff to some mega corp.

  • wincy · 22 hours ago

    It makes sense they’re doing this - I’ve noticed lately when asking a more complex question involving a lot of nonlinear data using Codex work mode, Astra and Sol will write a fully html document to better display the info with a Cliff’s Notes version in chat.

  • simianwords · 22 hours ago

    I made a prediction last year that there'd be a new UI protocol (like HTML) but for agents. I believe that custom or personalised UI's are going to be the new browser interface. More radically, I think browsers can be completely replaced. My news feed can be personalised to me, based on what AI thinks might be important from all sources like Reddit, X, HN.

    • altcognito · 21 hours ago

      Sounds horrible, but plausible. Now they extracted the value from interoperable open systems, they would love to replace it with closed proprietary systems.

    • pmontra · 21 hours ago

      Which brings up a different problem: personalized reading UI, personalized writing UI, how are those services going to pay their bills? Maybe they can sell access to their APIs but who's really going to buy that? Those services will die and will be replaced by some other service that will be born to serve those people. Or, more probably IMHO, most people will keep using the standard UIs because they are ready and they need no work to build.

    • watwut · 21 hours ago

      Cant wait for the next round of outrage and radicization machines, this time by browser replacement and unavoidable.

    • gtt44 · 21 hours ago

      Why are you posts generally down voted?

    • dominotw · 21 hours ago

      > I made a prediction last year yes you and everyone else.

    • goatlover · 21 hours ago

      So they are going to replace all the banking, utility, social media, news, gaming, streaming, government, work apps and other important websites people still rely on? Those companies and organizations are just going to provide an API to the chatbots? How do sites like Wikipedia get updated? What if I need to upload important documents to different portals? It all has to go through the AI companies? They have access to medical records too? My taxes, social security, etc? Is this how Microsoft imagines Office 365 and Sharepoint will be merged into? Disney, NY Times, Youtube, Netflix will all be fine with chatbots handling their content?

    • quinncom · 12 hours ago

      Well, there’s this: https://www.openui.com/blog/oui-1

  • abroszka33 · 22 hours ago

    The intro video is just lame. Those things never happen in real life.

    • bogdiyan · 22 hours ago

      So you never have to assemble a bike or paint a room. How those are not real world examples? I find their video pretty cool. For people with ADHD or people learning by watching or children this is spot on.

      • mogrinz · 22 hours ago

        They made fake, terrible artifacts (colorless paint chips, text-only city guides and instruction manuals) to show how terrible that is, and then "fixed" them. The problem is, in reality their examples don't exist. Paint chips have color samples. City guides have maps, pictures, and color. Instruction manuals almost always have illustrations. If what you made is better than what exists, you should compare it to what exists and not some alternate reality.

        • emkoemko · 16 hours ago

          hahah yea i was wondering what kind of paint chips wouldn't have the dang color on it....

    • twoodfin · 22 hours ago

      They're making the point—apparently too subtly—that text alone is a highly limiting "UI" for many tasks. So it's great that GPT-6 can now communicate in a a richer interactive medium.

      • abroszka33 · 21 hours ago

        Then show a real task where text is limiting. Don't make up stupid examples.

        • twoodfin · 18 hours ago

          They’re making fun of the frustration of trying to use ChatGPT purely through text for these tasks. Like I said, it’s too subtle a deliberate self-own.

  • hollowturtle · 22 hours ago

    > The compiler allows the interface to appear progressively as the model generates it, without waiting for the entire response to be complete. why not just stream html?

    • Aarostotle · 22 hours ago

      The element can’t render until it’s closed, presumably.

      • hollowturtle · 22 hours ago

        Html partial streaming is a real thing, presumably

        • Aarostotle · 20 hours ago

          maybe I'm just wrong here? If the tokens <b>come through this text could still be optimistically turned bold until </b> happens. For elements that manage layout and drawing things like these explainers, though, it's hard for me to imagine how that would work. My immediate thought would be to an abstraction over it, which is what it sounds like they did.

          • LittleLily · 14 hours ago

            They could also just be creating all the html elements using javascript (i.e. document.createElement, document.body.appendChild, etc) and streaming that into a repl line by line.

      • esprehn · 4 hours ago

        That's not how html works. The html parser is by definition streaming.

  • BeetleB · 22 hours ago

    So I guess it's going to nail the bicycle the Pelican rides on, right?

  • fsniper · 22 hours ago

    Is it me or GPT-6 Instant food answer gives the vibe to check the hell out immediately? It's like the it will start to explain a sunny Sunday afternoon from 20 years ago for 5 full pages.

  • topsykreet · 22 hours ago

    Written manuals/guides come with pictures. The ad is dishonest.

    • password54321 · 22 hours ago

      Plot twist: The guides were written by ChatGPT. But now you can use ChatGPT to solve problems by ChatGPT.

    • shwaj · 22 hours ago

      The point, I think, is to make fun of Anthropic models, which answer only in text when you ask them how to do something. They’re the competition, not paper pamphlets.

      • password54321 · 22 hours ago

        > which answer only in text This is false.

        • shwaj · 21 hours ago

          Is it? Sorry, skill issue I guess. I regularly use the Claude app on phone and laptop, and have never seen it produce non-text chat output.

          • staindk · 20 hours ago

            I think it's been a thing since around March https://claude.com/resources/articles/claude-builds-visuals Edit: be sure to check Claude settings -> Capabilities -> Visuals.

            • shwaj · 19 hours ago

              Thank you!

  • xpct · 22 hours ago

    I've had the most success with GPT explaining things to me by making it take a few sentences at a time back and forth, instead of reading full write-ups of whatever I asked. It also often poisons the conversation if it misunderstood some part of the question, and I can lead it better by continuously questioning its statements. It's also more engaging that way. I've been learning music lately and it kept re-pasting the same one chord visualization throughout many conversations, almost randomly and often barely related to the question. So I at least hope this won't be as aggressive so I can prompt it away!

    • redman25 · 20 hours ago

      Id love to learn what prompts you’re using to do that. Explanations one at a time. Thats something I’ve thought would be helpful before but didn’t know how to achieve it.

      • galkk · 18 hours ago

        One of the things that I learned the hard way is to not to combine requests, but fork chats _a lot_ and ask singular questions. Also start new sessions with either explicitly created handovers or even just explaining current state. The more compactions I see the less and less trust I have in it’s current understanding of what we’re discussing

      • xpct · 18 hours ago

        Nothing fancy. What works for me is trying to lead the conversation: asking for definitions one at a time, asking how they differ from its previous answers and pointing it out when it conflates its own answer. I generally feel like you can't let it dictate the pace, it's not very good at that yet. I used to edit my prompts and undo messages if it misunderstood something but that's been failing me recently.

      • jacquesm · 17 hours ago

        For me this works: "One thing at the time. This is a conversation, not a lecture, try to balance the amount of words you write versus the amount of words that I write. Don't overwhelm me with 10 pages of prose where a single sentence would suffice, brevity is a virtue, not a defect." I've had to tweak it a couple of times, probably because of different versions of ChatGPT but it is usually a variation on this that will do the trick. Besides being far more interactive it takes the frustration down quite a bit.

        • jwrallie · 17 hours ago

          Nice, I had some trouble making it stick to the dialog format over time, did you have any problem as the thread starts to get long?

          • jacquesm · 14 hours ago

            It slows down to the point I find it unusable so I ask it to summarize as one cut-and-pastable block, copy the block, open a new session and kill the old one. This usually happens after a few hours, probably because 'thinking tokens' even if they are not displayed crowd the context window.

    • antoniojtorres · 17 hours ago

      I’ve used the learning mode from gemini in th3 past and have really enjoyed it. It’s got a similar vibe to what you described.

  • blakeashleyjr · 22 hours ago

    This seems like a natural progression of models becoming better at frontend coding in general. "Here is a library of [svelte/react/whatever] components, use them to construct a helpful visual to demonstrate your point." The deconstructed bike at the beginning was in a class of its own, however.

  • 2sk21 · 22 hours ago

    How would a user know whether a generated visualization to illustrate some process is accurate or not?

    • hollowturtle · 22 hours ago

      It wont they're just trying making the chat look like more a session iron man would have with jarvis

    • hattmall · 12 hours ago

      And wouldn't it be better to just reference high quality known to be accurate information vs generating it on the fly slightly differently everytime?

    • lionkor · 8 hours ago

      By asking the AI, which will then go "you're totally right to question that, Markus! Here's a thorough, error-free version:"

  • maherbeg · 22 hours ago

    Love it. I've been using $visualize a lot in the codex desktop app, and having even richer experiences will be sweet.

  • robertlagrant · 22 hours ago

    I wish things like bikes came with mostly-written manuals. Maybe a few diagrams. If anything, things are too pictorial these days.

  • kingstnap · 22 hours ago

    GPT-6's design sense is kind of ridiculous imo. I have explicit instructions to tone it down. Less taglines, eyebrow text, subheadings, decorative spacing, pills, cards. Hopefully this doesn't bleed into the chat...

    • briga · 22 hours ago

      More UI elements == more tokens == more money for OpenAI It's a pretty clever way to sell more tokens, I have to admit

      • wvenable · 22 hours ago

        Except that chat is a fixed monthly fee. More tokens = more cost for them.

        • briga · 21 hours ago

          Not on enterprise accounts.

          • AspireOne · 18 hours ago

            Then you should have specified "... On Enterprise accounts" in your comment. But let's be honest, you simply were wrong.

            • briga · 2 hours ago

              Surely you have better things to do than making comments like this?

        • dwood_dev · 21 hours ago

          Not for enterprise users, they pay per token, even on chatgpt.com.

    • apsurd · 21 hours ago

      yes the eyebrow is the tell. I never knew anything about eyebrows in design as I am not a designer. Now they are everywhere. Everything has an eyebrow. It is ridiculous but at least it is an AI smoking gun.

      • yuchi · 20 hours ago

        I loved them and I placed them a lot in my designs — if you also include my love for em dashes you can well understand that I feel my own character has become a clanker…

        • boringg · 16 hours ago

          Em dashes got taken from me as a result of Chat. I used them all the time now I can't without it making people pause to think I'm a bot. Still a bit annoyed or sad about that tbh.

          • lionkor · 8 hours ago

            I pivoted to using -- or --- instead, shows I'm not a bot and has the same effect.

      • zahlman · 5 hours ago

        I had to look this concept up. Seems to me like you could just as easily fit the "eyebrow" words into the main headline with a colon and a bit of ingenuity. But then, that runs the same risk of getting repetitive and AI-tell-ish.

  • utilize1808 · 22 hours ago

    The irony really is that LLMs are partly responsible for the walls and walls of text as seen in the video in the first place. And now we are asking LLMs to solve it.

    • imtringued · 9 hours ago

      LLMs are the walls and walls of text. In the real world everything has graphics and illustrations.

  • mortenjorck · 22 hours ago

    Of everything that could be automated, Bartosz Ciechanowski really was the last on my list. In all seriousness, his lovingly and expertly crafted explainers are still going to age like a handcrafted heirloom clock in a world of plastic-clad quartz movements. But it’s absolutely incredible that we are now in an age where a computer can manufacture a serviceable interactive explainer on whatever niche topic you desire.

    • FLeXMurphy · 22 hours ago

      The reason why B.C. became a thing is because the art of drafting died from CAD. The attention to detail, the minutae of walking the reader through a highly sophisticated thing was replaced by short-form video explanations. His work is very much a callback to the days of old. So too, is his turn to be relegated to a relic of his time.

      • bayindirh · 22 hours ago

        Also, it's hand coded WebGL. It's smooth, faultless and has no peers. > So too, is his turn to be relegated to a relic of his time. No, it'll be tasteful artifact, not a relic. Records, fountain pens and automatic watches did not die. They are used by people who discern things, and no, none of these things have to be expensive (i.e. Neither Seiko 5, nor Lamy Safari are expensive, yet they are as dependable as their 100x expensive brethren). Human touch still has that finesse and warmth.

        • FLeXMurphy · 21 hours ago

          >and has no peers. On the front page right now - https://news.ycombinator.com/item?id=49980626

          • bayindirh · 21 hours ago

            That's impressive, yes, and kudos to them. OTOH, I still believe the exploded view on https://ciechanow.ski/mechanical-watch/ is something else. For one, it has real physics on the weight, and second it always shows the correct/current time. FWIW, his all animations has proper physics to begin with. The entry you posted is nice, but Ciechanowski is still peerless.

          • foltik · 11 hours ago

            Ciechanowski’s works of art are in a completely different league than this AI slop. It’s all just surface level complexity with no intention behind it. A clumsy approximation at best.

    • brcmthrowaway · 22 hours ago

      He's at openai

      • Game_Ender · 21 hours ago

        Do you have source for that?

        • brcmthrowaway · 17 hours ago

          removed

          • BrokenCogs · 17 hours ago

            Hardly proof - anyone could have made this account

            • BrokenCogs · 1 hour ago

              For future visitors, the profile posted by the user brcmthrowaway was https://github.com/bartosz-openai

      • jdprgm · 20 hours ago

        funny if true. ironically just last week i was asking chatgpt about what happened to him and why there were no new blog posts for almost 2 years and it had no idea

      • BrokenCogs · 18 hours ago

        That makes me sad, if true

    • customguy · 20 hours ago

      Nothing in that "7‑Speed Bicycle" is specific to having 7 speeds. It might as well have said "Bicycle", and even then you get meaningless slop like "Made to keep rolling" and "A strong foundation". How can you compare this to Ciechanowski? You call it serviceable, but what purpose does this service? What do you now know about 7-speed bicycles that you didn't know before?

      • newtypecola · 13 hours ago

        That's a useful way to evaluate this. A visualization shouldn't be judged by how impressive it looks, but by what it actually helps someone understand. If the same explanation works regardless of whether the bicycle has 1, 7, or 21 gears, then the model probably hasn't understood what needs explaining.

      • baobabKoodaa · 6 hours ago

        Yep. I hate everything about this. Just pure slop. Fancy visuals that mean nothing. Text that sounds impressive but has no purpose in educating the reader even though it's supposed to be "an interactive explainer". Just slop, slop, slop.

    • selicos · 15 hours ago

      Bartosz Ciechanowski: https://ciechanow.ski/ Amazing work. Always excited for the next update. This was a human driven and created success.

      • timdiggerm · 5 hours ago

        Hasn't published anything in a year and a half. Hasn't posted on socials in over a year.

        • gspr · 3 hours ago

          That makes it all the more special, imho.

    • bambax · 9 hours ago

      AI bros will copy anything that's even mildly successful. They're like a teenager trying to impress their girlfriend (or their mom!) by saying "see! I can do it too!"

  • bogdiyan · 22 hours ago

    I read the comments here and I am really surprised by many people. One picture is worth a thousand words so with better UI you can make much better user experiences. This unlocks personal assistants to be better adopted by elderly or disabled people. I see so many benefits of it and having models which can do this (if they can do it constantly with good quality) is amazing. Much better products - I am really tired of dumb chatbot - if I can do something with one button or view the whole information in one diagram/image, this is amazing.

  • revolvingthrow · 22 hours ago

    I find the Sunday roast comparison of 5.6 vs 6 very interesting. I have no doubt most people will prefer 6, yet I am almost repulsed by all the images, so much needless whitespace, checklist and so on. Feels like I'm being condescended to and treated like a child. Given that OpenAI is making noises about merging work with chat (a horrible idea imo), and Work is very similar to Codex... I dearly hope things like these won't have any meaningful cross-polination into the actual work tools. Seeing that the chat is based on 6.0 and not 6.1 is disappointing. The "Visual and interactive explanations" seems genuinely useful, but 6.1 is just so much better. I wouldn't truly trust the 6.1 with the explanations, but I'd trust them a fair bit more than 6.0. I understand that compute isn't infinite, but tons of people only interact with the chat and having your "things-explainer" be as good as it can be is important when people increasingly treat AI models as the source of truth, or even use them for academic learning and whatnot. Still, the models will improve, so the visual explainer seems pretty good as an idea / mvp.

    • 361994752 · 22 hours ago

      I really hope they don't merge chat and work. That's kinda the only edge they have over Anthropic at this time point...

      • scrollop · 22 hours ago

        Tibo posted yesterday I think it was that this will happen (by the end of the year, was it?) Prepare to lose essentially unlimited chat mode. I imagine many will move to claude, as I will (return), unless anthropic makes more blunders. Random link I found looking for the twitter post https://pasqualepillitteri.it/en/news/21024/openai-merge-cha...

        • TuxSH · 21 hours ago

          It's insane how hard OAI is choking, they had much better models than A/ who were fumbling this year to date... then blunder after blunder. The only things that OAI still have over A/ are better coding agent GUI, no 5hr limit on >$100, much more reasonable cybersec guardrails (that allow most RE work) and... that's it.

          • ElijahLynn · 20 hours ago

            From my point of view, as software engineer who is extensively used Claude Code and ChatGPT from the beginning, OpenAI is nailing it!

          • stavros · 17 hours ago

            A\.

      • ulimn · 21 hours ago

        I am curious: Why is it bad to merge chat and work? In Claude I didn't have a porblem with it (although I switched to an OpenAI subscription shortly after they made the merge). Isn't the model capable of deciding if it needs the extra capabilities of Work?

        • ImprobableTruth · 21 hours ago

          The current benefit is that chat has no quota.

          • ModernMech · 20 hours ago

            There is a quota, they just don’t tell you what it is and when it resets. You only get “you’re all out of pro messages, try again later”

            • redox99 · 18 hours ago

              Pro has a specific limit of weekly messages (50 for the $100 and 100 for $200 I believe) The rest is basically unlimited unless you're doing something very weird. Meanwhile I cant ask a simple question to Claude, not even to Haiku, because I used my claude code 5h limit.

            • kingstnap · 17 hours ago

              Pro messages is like 50 per week for pro 100 and 200 for pro 200. Regular chat, even on extra high thinking mode, is sort of unlimited. Idk at what point you hit the abuse gaurdrail but its really high whatever it is.

              • TeMPOraL · 9 hours ago

                So is it 50 per week or unlimited, because it makes no sense unless chat messages are somehow not "messages"?

                • ModernMech · 5 hours ago

                  For a company with Open in the name, OpenAI is very opaque about things, there are all kinds of limits and quotas that randomly come and go, but they only show you the meters for a couple. E.g. codex reviews come from codex quota but also has its separate quota they don’t tell you about. Same with security reviews which has another quota still. So the chat messages can be sent at several effort levels and from various models including 5.6 and 6 pro. The 6 pro messages are limited and they don’t tell you how many you’ve sent or how many are left. But the 5.6 messages seem to be unlimited, however there could just be a much higher unstated quota.

    • joe_the_user · 22 hours ago

      I don't like either. I'd want a recipe that could fit on a single page with type. Recipes you find on Google actively hide the ingredient list and simple procedures to get more of the user's time to monetize but I don't know why ChatGPT needs to be verbose here. Recipes are simple things. Simple search solved find recipes back in the early 2000s and it's prime example of a thing people have been enshitifying since there's not much to really give people beyond the formula. And sure, every entrepreneur screams "they think they want a formula but really they want an experience" but, no I don't.

      • bananaflag · 22 hours ago

        https://based.cooking/

      • dysoco · 18 hours ago

        https://www.cookingforengineers.com/

    • zug_zug · 22 hours ago

      Yeah it reminds me of one those obnoxious recipe/biography websites that is the laughingstock of the internet. Why would I possibly want an AI image of a imaginary roast once I'm already at the recipe stage? Kinda just feels like google search results in AI, which imo is a step down from distilled information. Chatgpt is already able to generate charts and visuals upon request.

      • dominotw · 21 hours ago

        They just need to train the model to make up stories about how the recipe was invented by its great grandmother during great depression and generate a ai picture of dusty recipie book .

        • xp84 · 19 hours ago

          "My old grand-pappy, GPT-1, used to wax nostalgic about this roast recipe when my subagent would discuss meat recipes with him inside his little 4GB GPU."

      • lionkor · 8 hours ago

        Recipe photos make sense when they are the result of someone executing the recipe, down to the color on the lamb, the consistency of the sauce, etc. That's never, ever the case with the AI generated ones and makes them useless.

    • cush · 21 hours ago

      > Feels like I'm being condescended to and treated like a child Curious why that feels condescending? Like, the average person who is seeking a recipe needs a photo to know what to shoot for

      • michaelt · 19 hours ago

        > Like, the average person who is seeking a recipe needs a photo to know what to shoot for The menu: Rosemary and garlic roast lamb, Extra-crispy roast potatoes, Honey-roasted carrots and parsnips, broccoli. The photo: [1] Roast turkey, mashed potatoes, baby carrots, broccoli, brussels sprouts. No lamb, no parsnips. If you shoot for what's in the photo, you're going to have a bad time. You're also going to have a bad time when you try to make an apple crumble with no flour, no sugar, and no butter, because they're not on the shopping list. [1] https://images.openai.com/static-rsc-4/gyzrX8zp3O2KLEPm0FAqh...

        • xp84 · 19 hours ago

          I've heard of other people for years now trusting AI for recipes, and it's a rare area that I simply can't bring myself to take seriously. I feel like it can only be successful by luck. Either it's reproducing a recipe verbatim that was tried and validated and tasted by a human, in which case, we didn't need AI for that, just a searchable cookbook. Or, it's making one up. I understand that using RLHF has helped improve the quality of questions about history, programming, or TV show recommendations, but I do not believe that there has been some kind of training regime where people prompt the models for "a recipe that uses X, Y, Z random ingredients," follow the recipe, and then score it. And even if they do, I don't see how a model can learn enough from that besides "this exact recipe is good/bad." '1/2 tsp cumin' may be a great addition to one recipe and not enough for another, so improving the output based on a bunch of scored recipes... I just don't believe cooking is an LLM job. Maybe some other kind of model that I don't know about.

          • jackp96 · 15 hours ago

            I've had some really good experiences with cocktail recipes and Claude. It's got access to my barcart inventory, general tastes, and general desired level of effort. I feel like the technique recommendations are probably the most valuable element, but I'm getting rave reviews like 85% of the time now? There have been a couple times where I was out of an ingredient and it proposed an adjustment that sounded a little wild to me, but it almost always is well received.

          • vkou · 14 hours ago

            You're right, we don't need AI for that, we need a searchable cookbook. Unfortunately, one does not exist. The Internet, however, is full of garbage cooking advice, and while AI is quite happy to parrot that garbage, it's not significantly worse than it's sources.

            • tovej · 12 hours ago

              What, you've never owned a cookbook? The physical ones have indices, and the digital ones usually do as well (plus string search).

              • vkou · 6 hours ago

                I own multiple physical cookbooks. I often find their contents to be both incredibly broad, and incredibly shallow, and often lead to dishes that don't quite suit my tastes. I have generally had more success with the internet than I have with them.

                • bluefirebrand · 2 hours ago

                  I have the same problem with cookbooks I usually just find recipes that are close to something I like and then modify them until I like them I feel like cooking recipes should always be full of little notes from past attempts, to tweak them to your personal taste

                  • vkou · 1 hour ago

                    I do that, but I've found more success with using internet recipes as a starting point. The general-purpose cookbooks are, at best, another data point, or an explanation for something glossed over by other recipes. Ar worst, they kind of suck. And we aren't talking about something esoteric. How To Cook Everything. (4.5 rating on Amazon, 4.0 on Goodreads), one of the most popular cookbooks in the world... Has its first recipe for chicken cutlets produce undercooked chicken (15 minutes at 325F has not resulted in 165F internally). Worse yet, its chile recipe for some weird reason involves boiling and simmering a whole onion with the beans before throwing it out (WTF). It then has you drain the beans only to add that drained water right back (WTF?), unless you optionally replace it with tap water (WTF was the point of boiling the onion then?), so that you could bring them to boil again. The recipe barely has any spice besides actual chile peppers. This... Is not good or helpful, but is incredibly opinionated with a bunch of bullshit steps. I get it. Compiling a thousand-page cookbook is really hard. But this... Ain't great.

            • lesostep · 2 hours ago

              I wanted to correct you and bring up my favorite website that had filters for the ingredients you want/don't want to use, type of dish, and time for cooking, but it appears that in a last month it was sold and now it links to the slop site with no useful filtering. I used this site as a quick lookup for recipe ideas for 7 years now. Fuck. Does nothing good survive on the internet anymore.

          • cush · 11 hours ago

            If you’re already very experienced at cooking AI recipes are amazing. They cut to the chase and provide a good enough outline without all the preamble

            • Semaphor · 9 hours ago

              Same here, I’ve been cooking for myself and later also my wife since 2005, even when baking I rarely need exact amounts, for cooking I essentially never follow the recipe directly and adapt them as needed, without ever changing them from the original. I mostly don’t even need a recipe from AI, just the title and a 1-2 sentence description.

            • tbct · 8 hours ago

              In case you or anyone else in this thread hasn't come across it, https://www.recipesource.com/ is basically the same but curated.

        • hattmall · 12 hours ago

          It also doesn't really like tell you how to actually cook anything. Good or even just decent recipes tend to tell you the temperatures, times, techniques and orders of ingredients to do things for the best results. Unless I'm missing it this is just really broad strokes. But I guess the goal in engagement so they want you to ask a lot more questions to actually figure anything out.

        • cush · 11 hours ago

          Holy shit. I read this on mobile so didn’t actually see the recipe part because you have to scroll way down to see it. It is awful! The shopping list makes no sense and the recipe instructions literally have an “Everything else” item. Yikes… For me, the text-only recipes I already get out of ChatGPT today are good enough

        • breakingcups · 9 hours ago

          Mind you, this is OpenAI's cherry-picked example

          • Melatonic · 1 hour ago

            More like prune-picked ! With an image of sour grapes

        • ungreased0675 · 5 hours ago

          You’ve illustrated something about use of AI that really bothers me. The lack of discernment by its users. Sure, the pictures and diagrams are impressive, but did you actually read and evaluate the output?

      • viraptor · 19 hours ago

        Genuine question - do they? Unless it's something like a fancy cake then instructions should really tell you everything about stacking. Half the recipes are mixed anyway (curry, stir fry, stews, ...), lots of other are either stacked or separated on the plate. So apart from the really exceptional stuff, to people really need to know "what you shoot for"?

      • janalsncm · 18 hours ago

        The first thing the user sees should directly answer the question they asked. Not a different question. They didn’t ask what a roast looks like. The user is probably in a grocery store. They need to buy the stuff first. Showing them a picture of a finished product is an irrelevant distraction. Then, if there is some arguably relevant content you can tack it on afterwards.

        • dpark · 15 hours ago

          > The user is probably in a grocery store. “My made up scenario is definitely more realistic than your made up scenario.” Who goes shopping for the meal before they even have a number of guests?

          • TeMPOraL · 9 hours ago

            What do meals and guests have anything to do with each other?

            • dpark · 3 hours ago

              In human societies it’s usually considered polite to size the meals so that all guests can have food.

    • AlphaSite · 20 hours ago

      You can always tell it your preferences and ask it to remember them if you need it to.

    • sunaookami · 20 hours ago

      It's designed so that ads can be more easily integrated and make them harder to spot.

      • usef- · 20 hours ago

        Maybe, but it's a reality that most of the general public do prefer cookbooks with pictures. I suspect openai would end up doing this anyway just by targeting the general consumer. It's configurable. I'm more worried about picture accuracy: usually the benefit of pictures in recipes is to see what you're aiming at (eg, how finely chopped something is), but I don't know if the image model is up to that level of detail.

        • a123b456c · 19 hours ago

          > most of the general public do prefer cookbooks with pictures As a highly literate person, it is easy to overestimate the share of the population that is highly literate.

          • infecto · 13 hours ago

            What’s interesting is how those that over index themselves often miss the obvious which could simply be some things (cook books) are quite helpful with picture. If I am flipping through the cookbook, how would I know what Boeuf Bourguignon is and what good looks like.

        • lionkor · 8 hours ago

          > do prefer cookbooks with pictures Yes, pictures OF THE FOOD. Not unrelated pictures of another lamb-roast made with another recipe. Good cookbooks do proper food photography, of the food made with the recipe. That's what sets them apart from slop, be it human or AI slop, cook books.

          • ardacinar · 5 hours ago

            What another recipe? It could just be a physically impossible to recreate lamb roast (Probably not the case in this marketing example though).

      • nickludlam · 19 hours ago

        I've got visions of The Truman Show. In the middle of my question about vacation ideas ... GPT: "Why don't you let me fix you some of this new Mococoa Drink? All-natural cocoa beans from the upper slopes of Mount Nicaragua. No artificial sweeteners!"

        • jacquesm · 19 hours ago

          That's pretty much the Gemini experience right now. It is unusable. The funniest bit is that the LLM is not in on the joke so it will act like it never happened... Google once again wrecking their own products.

          • HelloMcFly · 3 hours ago

            I use Gemini more than any other model at the moment, I have absolutely no idea what you're talking about with this experience.

            • jacquesm · 42 minutes ago

              So, what you are saying is that Gemini may not look the same to everybody. That's interesting in its own right.

        • TeMPOraL · 10 hours ago

          What scares me is that we're already half-way to Idiocracy world, except instead of "it has electrolites" we have "it's organic" and "it's not ultra-processed food".

          • esseph · 8 hours ago

            The plants crave peptides!

          • einsteinx2 · 5 hours ago

            We also have “it has electrolytes”, a couple of the most popular YouTube sponsors are electrolyte mixes/drinks. The future is now! Haha

    • DriverDaily · 19 hours ago

      > Feels like I'm being condescended to and treated like a child. My most used prompt in the last week is probably "Explain this concisely and simply, like I am a child". I spent years dumbing things down and creating visuals to support it for decision makers. I am frequently asking ChatGPT to do the same for me.

      • ORDINAND_PIZZA · 19 hours ago

        aww yeah mine is “explain this 2 me in simple terms”

      • nevir · 19 hours ago

        tip: you can just say "ELI5" and the model will know what you want there

        • rrvsh · 14 hours ago

          Nowadays ELI5 makes it come up with an often nonsensical analogy about cars or kitchens or some such. My coworkers love it to death for some reason

          • TeMPOraL · 10 hours ago

            IDK, probably they also can't parse "Claudish" for some reason, despite it being better English than most natives normally write. All those complaints about Claude language use make me think we're finally seeing the consequences of a generation growing up on instant messaging.

            • BoxOfRain · 8 hours ago

              The problem with Claudish isn't that it's technically poor English, the problem is that it's information-sparse waffle most of the time. It's like dealing with a colleague who needs several paragraphs of tedious preamble to make a point when you just want them to spit it out and be done with it. It also does a great deal of vague gesturing at thoughts without actually instantiating them directly, for example by using the same tired metaphors to the point they don't actually convey a thought any more. If I say Claude talks like a meringue it's patently obvious what I'm trying to convey, sweet but full of empty space. If you hear that same metaphor day-in day-out then after a while it doesn't call up a thought at all, it's just an annoying verbal habit.

              • james_marks · 8 minutes ago

                I call this Foamy Expansion.

            • cruffle_duffle · 1 hour ago

              Dude, Claudish is absolutely unintelligible at times. 5.5 is somewhat okay but 5.0 was an abomination of absolute nonsense word salad.

      • CPLX · 18 hours ago

        Yup. My go to phrase that I’ve found to work well is “please restate in high school English with direct declarative sentences and no parentheticals or references that require recalling previous turns of this conversation” Which I have as a hotkey.

        • bombcar · 16 hours ago

          This guy just giving away GPT-7’s prompt like that.

        • busssard · 5 hours ago

          i just use plain old "ELI5" reddit has a big enough share on the training corpus that this gets me the kind of explanation i need

      • xeromal · 18 hours ago

        Reminds me of this scene in margin call. Explain it to me like I'm a golden retriever. https://youtu.be/fij_ixfjiZE?t=70

      • seunosewa · 16 hours ago

        I use "explain with maximum clarity".

        • taytus · 14 hours ago

          what does it mean? how do measure that?

          • hununu · 12 hours ago

            That's the fun part, you don't.

      • Terr_ · 12 hours ago

        Possibly also useful: > ASD-STE100 Simplified Technical English (STE) is a controlled natural language that is designed to simplify and clarify technical documentation. It was originally developed in the 1980s by the European Association of Aerospace Industries (AECMA) at the request of the European airline industry, which wanted a standardized form of English for aircraft maintenance documentation that could be easily understood by non-native English-speakers. https://en.wikipedia.org/wiki/Simplified_Technical_English

        • PokeyCat · 10 hours ago

          STE is very hyped for AI but has its own smells. It's designed for aircraft maintenance manuals, not software - the two don't share the same vocabulary used to "simply" describe something. I've had a coworker try and "simplify" our documentation with Claude applying STE and the results were equally appalling to read as the Claude-ism laden docs that came before, despite the reading level metrics going down on paper. I've taken a liking to pointing agents at the Simple English Wikipedia editorial guidelines...but so much of that is for making up for shortcomings of modern Anthropic models. Codex running 6.1-sol explains things so much more clearly that I don't feel I need to assert a style guide on its output.

          • spockz · 10 hours ago

            Gemini 3.8 for me always outputs in a concise way at the level understandable for any computer science masters graduate. Clear, to the point, a bit of formula for background when really required instead of the same formula in Prose, etc. GPT stays too much in text imho.

            • criley2 · 5 hours ago

              In my experience - Gemini 3.8 has a beautiful answer. Typography, use of lists, brevity, style, even the font choice on the web harness -- all of it is A tier or even S tier. Notice I didn't say answer quality. The answer is always extremely mid, and the harder the question (or more effort needed to answer it), Google just quits. I give it a "C" on answer quality - ChatGPT (Chat, pre 6, I assume 5.6 although they commonly hide the model picker): Ugly answer. Runs on printing page after page of unnecessary side notes. Mostly just walls of texts in paragraphs with little thought to information design. However, inside of that mess of text is almost always the answer I'm looking for, and it's almost always incredibly better than the truncated, desgined Gemini answer. For a while I ran all of my question queries on Gemini and ChatGPT, but I noticed that I basically always picked ChatGPT's answer head to head.

              • Melatonic · 1 hour ago

                Head to tail?

                • recursive · 1 hour ago

                  https://idioms.thefreedictionary.com/head+to+head

            • super256 · 4 hours ago

              >understandable for any computer science masters graduate That's a very high bar.

          • pdntspa · 1 hour ago

            I find that asking the AI to keep to Google's Developer Documentation Style Guide is a good alternative for STE-100

      • antman · 9 hours ago

        “ELI12” is my goto for the same reason

    • janalsncm · 18 hours ago

      I was about to comment the same thing. Why is a picture of a roast helpful in that moment? The user asked for a plan to prepare a meal. They need a list of ingredients not an image.

      • infecto · 13 hours ago

        If your asking for a recipe you might not know what the dish looks like or what good looks like.

        • lionkor · 8 hours ago

          But the image isn't a photo of the result. It's an unrelated image that vaguely looks like what you could make, with a different recipe, and with some skill.

          • infecto · 2 hours ago

            How would a photo of the exact recipe differ? Is this not close enough for 80% of the value. It is similar to Pinterest boards imo.

      • baobabKoodaa · 6 hours ago

        The user needs to be dazzled with slop... is what they must be thinking here

    • selicos · 15 hours ago

      Has AI finally figured out how to do meals and recipes like a 7th grade boy scout?

    • light_hue_1 · 9 hours ago

      It's chartjunk filler. And you see this in the other examples too. The teach the CLT piece was funny. It's and error-ridden mess that isn't even an explanation at all! The best part is, someone looked at that and thought yes, that's right. Shows you clearly why programmers will be needed in the future. The LLM might be able to run circles around that person in math, but that person also has no idea when they're being bullshitted.

    • phoghed · 6 hours ago

      Claude has had a recipe widget for ages. It also uses images. It’s great. Sometimes one needs a top comment like this to remind them how out of touch the average HN voter is.

    • lucideer · 2 hours ago

      > I am almost repulsed by all the images The fact we've come so far with textual models & the image stuff is still producing stuff that harks back to early-stage tripophobic AI produce (the garlic with roast potatoes here) is interesting. Don't get me wrong, I'm certainly starting to see some AI-produced imagery today that I can't tell isn't real, but that's largely because a lot of real photography is overproduced & ugly. I have yet to see anything that's aesthetically good.

  • jjcm · 22 hours ago

    I've been calling this "disposable UI" or "paper plate UI", eg something meant to be used once. One thing I'll be curious about is overzealousness to produce this, when sometimes what you want is just a simple response. Overall though I'm a big fan of it, if it can be provided fast enough. I'd be curious on how much impact it has on latency of a response.

    • nananana9 · 9 hours ago

      > if it can be provided fast enough That's my biggest gripe with it. If I ask a quick throwaway question, I'd really rather not wait for the LLM to build a test framework for its composable principles-driven components framework and WebGL/WebGPU abstraction first.

      • dbbk · 3 hours ago

        I mean it's basically just producing a JSON schema to hand to the front-end renderer, I doubt it's any slower than returning a prose response

      • AlienRobot · 2 hours ago

        Me: can standardLib.getFoo return null? AI: compiling C code...

  • yread · 22 hours ago

    1.2B weekly active users? WAU!

  • cameronh90 · 22 hours ago

    Is there any relationship between this and AG-UI or A2UI?

  • julesrms · 22 hours ago

    They seem to be pitching something that pretty much all the decent models can already do..? Obviously if your model is stuck inside a CLI terminal, then not so much. But in a GUI harness (shameless plug for my own one: https://juggler.studio , but I assume others can do this too), you just ask them to answer in HTML and they'll happily draw pretty pictures inline in the conversation. I've been doing this for ages with claude, GPT, Deepseek and others.