Rendered at 22:27:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
firasd 1 days ago [-]
I think it’s a misread to think the default ChatGPT model switching to GPT 5.6 Luna is some sort of desperation move. Keep in mind that Claude .ai never had this extreme stratification between the frontier models and the free tier (Sonnet is available to free users with rate limits).
So 5.6 Luna is just their next version of what they used to call 5.5 instant tier
And 5.x instant models were never much to write home about anyway so the default ChatGPT free model hasn’t been particularly distinctive since 4o
throwaway473825 9 hours ago [-]
I think the main competitor to OpenAI's free tier is Google Search's AI Overview powered by Gemini 3.5 Flash-Lite. OpenAI's free tier needs to be notably better than that for people to bother. There's also Gemini 3.6 Flash with a generous free tier.
stephbook 6 hours ago [-]
I'm getting all my questions (3-10 each evening) answered by "Gemini Pro Extended" in the free tier, sometimes with PDFs or photos attached. I wouldn't bother with anything "free" below that, unless I'm agentic coding or something.
dotancohen 5 hours ago [-]
Are you certain they are the correct answers? No other AI tool halucinations so confidently for me as does Gemini. And I get them across so many topics.
lxgr 3 hours ago [-]
> So 5.6 Luna is just their next version of what they used to call 5.5 instant tier
I wonder if this means that Luna is more, uhm, "free tier like" in its responses? The "Instant"/"Chat" models have a pretty particular vibe.
samsolomon 10 hours ago [-]
Honestly, given the price 5.6 Luna is pretty fantastic. I replaced GPT 4.1 mini in several projects.
It's also 50% cheaper if you only need it to run sometime in the next 24 hours.
MarkMarine 4 hours ago [-]
Flex tier is just as cheap but you don't need to deal with batching or waiting. I see it completing in about ~3x-10x the latency of normal calls, so in seconds. You can always route back to regular if you're getting 429'd on flex too.
Flex is also easier to get caching to work, there is a little futzing around with OAI's implicit caching but if you do the upfront work you can get haiku quality responses with caching in close to realtime for 1/7th the cost and you don't need to design a polling loop to check for batch completions
electroglyph 1 hours ago [-]
yep, agreed. little bit smarter than deepseek 4 flash, but a little bit worse at tool calls.
freakynit 19 hours ago [-]
Claude's rate limits are absolutely shit for free users. 3 chats and its gone for 24 hours.
drnick1 7 hours ago [-]
Yes, but you can create multiple accounts. It has to be inconvenient enough or no one would pay for the higher tiers.
robomartin 4 hours ago [-]
It's one thing to be a free user...quite another to be an entitled free user.
This is precisely why I do not offer anything for free. I've been there and done it back in the iPhone 3 era with a dozen apps. Learned my lesson quickly. Never again. It was a nightmare. If the product isn't good enough for someone to pay for it (that could be direct user-paid or advertiser supported) it should not exist.
asadotzler 2 hours ago [-]
If your free tier is marketed as a free tier but really it's just "your first hit is free" like your crack dealer's marketing, then it's a fraud on the user, a scam to get you hooked and paying. I am not an entitled user for thinking a free tier designed to provide half the answer and make me pay for the second half is a scam from an evil corporation and to call that out.
Also, it is intended to be supported by ads, so it fits your ideal product description quite well despite being a scam.
claudiug 14 hours ago [-]
I dont see that, and I have a lot of chats, long messages, and is not 3 chats.
vezycash 13 hours ago [-]
It has limits when there's an attachment in that chat session. Switch to a new chat and you're good to go
ilaksh 1 days ago [-]
> Our mission is to ensure that artificial general intelligence benefits all of humanity. We’re introducing updates to ChatGPT that improve everyday conversations while expanding access for Free users.
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.
kkoncevicius 1 days ago [-]
If we could show the current models to someone like Alan Turing, I am sure he would conclude that we have AGI.
throwatdem12311 12 hours ago [-]
You say that they pass the Turing test yet every post on HN complains about the way LLMs write so clearly they haven’t passed it yet because we can still tell it’s a bot.
pinkgolem 12 hours ago [-]
You can also change the writing style with a prompt quit easily, if the same person would publish millions of articles, we would also recognize them.
throwatdem12311 11 hours ago [-]
The Turing Test is about being able to figure out if “someone” is a computer during a short conversation, not “millions of articles”.
LLMs still live in the uncanny valley and can be sussed out immediately.
For example, I’m extremely annoyed by the fact that offshore developers respond to me almost exclusively using text generated by Claude. You can tell immediately because they use overly descriptive techno word salad that no normal human uses unless they are trying to be ultra specific for a scientific paper - and even then it’s still too much for a real person.
pinkgolem 3 hours ago [-]
My argument is, you recognize the style because you know it..
5 years ago, you would not have been able to determine that this was the case, and just have assumed it's a know it all character.
Tell your offshore developers to use caveman or so
PunchTornado 11 hours ago [-]
I don't understand. do you think that people can't figure it out if they speak with llms or not from a few turns?
quantumspandex 11 hours ago [-]
We don't train current LLMs to mimic the average human's writing style. We train it to be smart, helpful, and knowledgable, too knowledgeable for a human. We can easily train a LLM to pass the Turing test if we wanted to, but then it would just sound dumb or biased.
virgilp 9 hours ago [-]
> We can easily train a LLM to pass the Turing test if we wanted to, but then it would just sound dumb or biased.
Interesting idea. "This is not dumb and biased enough, probably not a human".
throwatdem12311 11 hours ago [-]
> but it would just sound dumb or biased
We have one of those: Grok.
hobofan 5 hours ago [-]
Yes, and it's indistinguishable from the average X user.
throwatdem12311 8 minutes ago [-]
Bold of you to assume they’re actual users.
guelo 9 hours ago [-]
There is tons of money trying to get customer support bots to sound human but they fool nobody.
hk__2 9 hours ago [-]
I think we’re just adapting. LLM felt kind of magical at first, and now we’re all experts in detecting AI slope.
3 hours ago [-]
chippiewill 12 hours ago [-]
Turing was a helluva smart guy, but he doesn't have the benefit of hindsight.
stymaar 1 days ago [-]
And after half an hour using it he'd just admit that his test was way too simple as these models are still way too dumb
klibertp 24 hours ago [-]
I think of the Turing test as one of the starting lines, along with image recognition ("a summer break project for a group of grad students" resisted being solved for decades).
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
jaccola 14 hours ago [-]
I don’t think he would. The Turing Test as originally formulated is a bit ambiguous but by most non-incentivised interpretations LLMs do not pass it.
In the original test the evaluator knew one was a machine and one a human and could have conversations of arbitrary length.
stavros 21 hours ago [-]
How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).
txzl 15 hours ago [-]
I just asked Fable 5 max to create a 2d game about caterpillar climbing a tree and eating fruits.
The game looks good - animations, 8-bit aesthetics, procedural tree branching, but the tree's branches are dead ends. You can't go back once you started climbing a branch.
Yes, LLM doesn't have a reliable way to test it's game yet. All the screenshots, and playwright tests will never be enough to test even a simple game.
But can we call a machine doing such mistakes a general intelligence? It has no embodied intelligence. No way to experience time the way we do. All it has is text. Yes, they can have images, sound too, but no big models (at least those we are supposed to use for coding) currently are native with video as far as I am concerned. And I am not sure that just video without embodied experience is enough to understand the world the way humans do.
Of course we can get incredible results from machines that have a very different experience of the world than we do. But is this a general intelligence? I guess "general" is supposed to mean being able to do everything any human can do (minus the skills requiring a body)?
ipsod 10 hours ago [-]
Current Gemini Flash models can take video input. They're not hyper-specialized coders, but they're better than the competition on many tasks. They seem to be better with spacial reasoning, as well - they are the best choice for OpenSCAD, for example.
stymaar 15 hours ago [-]
I think it's Karpathy who coined the term “jagged intelligence”. LLMs are both extraordinary smart in domain they have been explicitly trained on (like Math) and positively dumb on things they haven't.
Eggpants 9 hours ago [-]
Would you also consider a database of questions and answers smart? LLM are basically lossy text compression databases with a clever query method. Useful for sure but it’s not thinking, it’s recall.
Yes there are a few sites with IQ test benchmarks. The frontier models come up around 130 or 140 depending on which model/test.
nearbuy 15 hours ago [-]
It also looks like they're saturating the test, with one LLM hitting the maximum possible score. (https://www.trackingai.org/home)
The test wasn't made to accurately measure IQs that high.
stavros 21 hours ago [-]
That tracks, I wonder what the people who say that the models are too dumb expect to see. Miracles?
stymaar 15 hours ago [-]
Not making mistakes my four years old would not.
They are definitely smart enough to be useful, but dumb enough in their weak spots not to deserve the "general intelligence" qualifier.
qup 7 hours ago [-]
Well your generally intelligent four year old can't meet the standard you just made, by definition.
hluska 21 hours ago [-]
Dumb is too simplistic, but they sure don’t pass for being human. That would be the intent of a Turing test… not IQ.
ozgung 15 hours ago [-]
I wonder if it’s a version of Dunning-Kruger effect to call AI models dumb. I haven’t seen a “dumber than me” model since years. Also the smartest people known in the world use them in their fields so I don’t know what is meant by a “too dumb” model.
hnbad 13 hours ago [-]
You need to be smarter (or rather: more knowledgeable in the problem domain) than the model to be able to use it efficiently. Hallucinations are still a problem occasionally but a bigger one is failure of imagination. Even Claude Fable lacks a holistic understanding of many domains it wasn't obviously trained on. The biggest problem with AI (if we assert that LLMs can be the basis of AI) is that these models will make mistakes that exist in an entirely different category of the kind of mistakes humans will make.
As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.
kkoncevicius 9 hours ago [-]
> much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied.
This sounds interesting on it's own. I would be curious to hear more if you are willing to share.
orbital-decay 18 hours ago [-]
It doesn't fit their own definition of AGI, although it's inevitably vague [0]:
>artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work
I think that's the crux. As soon as something turns into a commodity, it will be priced as a commodity. It stops being economically valuable and can only be economically viable at scale.
Take mathematical calculations, styrofoam, LCD screens, ice cubes, embroidery (try searching for "computer work blouses"!), texts, navigation. The thing itself costs nothing, the service and theater around it becomes everything.
pure_magic 4 hours ago [-]
I don't think this means they believe ChatGPT is AGI.
whazor 14 hours ago [-]
My ChatGPT gives me lots of “artificial general intelligence” on a daily basis that I don’t possess.
kubb 1 days ago [-]
> This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Well, the models are smart enough to point out why this is wrong.
lostmsu 20 hours ago [-]
They were trained to do that specifically, don't you think?
BedVibe_Studios 9 hours ago [-]
I found answers from free tier models to be harmful, so helping humans in general isn't what I actually noticed here.
unreal6 8 hours ago [-]
Do you have an example? What type of harm?
braebo 3 hours ago [-]
Fable is AGI. All other models are dumb as bricks in comparison.
khalic 14 hours ago [-]
ok so first the industry takes the AI term, uses it with a very vague resemblance to what it used to mean, for marketing purpose, then they latch on to the AGI term to talk about what people used to consider AI. Now there is no sign of actual AGI happening any time soon, so we're going to reinterpret AGI to mean something diminutive like a chatbot?
Don't you see a problem here? Terms are used to describe the world and need a semblance of stability so we don't end up in a race to the bottom just so investors can feel good.
tristanj 13 hours ago [-]
> Now there is no sign of actual AGI happening any time soon
Do you genuinely hold this position, or do you not realize how far the goalposts have shifted?
In 2022, prominent AI critic Gary Marcus offered to bet $100,000 that we wouldn't have AGI by 2029. https://garymarcus.substack.com/p/dear-elon-musk-here-are-fi... Because the definition of AGI is unclear, he defined that AGI would be achieved if an AI model could do THREE of the five following tasks:
- In 2029, AI will not be able to watch a movie and tell you accurately what is going on (what I called the comprehension challenge in The New Yorker, in 2014). Who are the characters? What are their conflicts and motivations? etc.
- In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI.
- In 2029, AI will not be able to work as a competent cook in an arbitrary kitchen (extending Steve Wozniak’s cup of coffee benchmark).
- In 2029, AI will not be able to reliably construct bug-free code of more than 10,000 lines from natural language specification or by interactions with a non-expert user. [Gluing together code from existing libraries doesn’t count.]
- In 2029, AI will not be able to take arbitrary proofs from the mathematical literature written in natural language and convert them into a symbolic form suitable for symbolic verification.
Today's AI models can do FOUR of these five.
Using 2022 goalposts, we already have AGI. We blew past these goalposts months ago, and nobody noticed.
khalic 12 hours ago [-]
That’s his definition and in no way a universal one.
In the meantime, it’s still very easy to differentiate between an AI and a human in a chat. You just need to know the quirks of these systems. Like counting letters, hitting the safeguards, etc.
So call me when one of them can pass the Turing test against me and then we can talk about AGI
tristanj 2 hours ago [-]
> So call me when one of them can pass the Turing test against me and then we can talk about AGI
Ring ring I'm calling you right now. We blew past the Turing test goalpost over a year ago, using 2024 models.
GPT-4.5 passed the Turing Test with a 73% human rating, outscoring actual human subjects. That is, human evaluators considered the AI more human than an actual human, 73% of the time. LLaMa-3.1 was judged to be a human 56% of the time.
The 'strawberry' test was fixed years ago with the invention of CoT; models only fail that test today when thinking is disabled.
pixl97 9 hours ago [-]
There is no accepted definition of intelligence that is usable for classification of intelligence across the sciences.
On top of that for your strong feelings, you don't have the conviction to write down a strong definition of intelligence yourself, which allows you to accelerate the goal posts up to light speed. The fun thing about writing out a formal definition is suddenly almost everything or almost nothing, including a lot of humans, has intelligence.
Not basing intelligence on your feelings of the moment makes it a hard thing to define across everything intelligence applies to.
khalic 9 hours ago [-]
[dead]
andai 1 days ago [-]
>It also needs to be differentiated from ASI with godlike powers many times greater than human.
Wasn't there a thread the other day about how very few humans on earth can understand the new math proofs?
Although "with sufficient study" vs "not even with unlimited study" are probably worth distinguishing there.
Alifatisk 16 hours ago [-]
> Wasn't there a thread the other day about how very few humans on earth can understand the new math proofs?
Do you perhaps have a link to it? Tried to search around but couldn’t find it. Would love to read it
ilaksh 1 days ago [-]
Which brings up another point in that we should have a term that distinguished between super intelligence in some capacity and godlike super intelligence.
ozgung 15 hours ago [-]
It’s interesting that they haven’t declared AGI yet, even as a PR stunt. It can be like a “pre-revenue” tactic. They are pre-AGI so investors can still pour money.
tristanj 13 hours ago [-]
Public models are months behind what the labs have internally, and I believe OAI's upcoming model is so advanced they consider it AGI.
OAI is not going to announce anything until their next model officially releases (allegedly later this month).
jonaustin 4 hours ago [-]
> This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
> Which I think is a fair interpretation of the term
You can't be serious... Case in point - I just asked GPT Sol High to give me a weekly update of local ai changes.
Here's it's first update, which is complete and utter garbage; i.e. it's a lot of words that says absolutely nothing.
That's just a random word generator; AGI? Not even remotely in the ballpark.
Well your comment adds no value and is just offensive. Why don’t you argue a counterpoint?
aryehof 18 hours ago [-]
For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show.
It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?
dannyw 10 hours ago [-]
A lot of users are very latency sensitive. The median ChatGPT.com user is not like the HN community; many of my friends use ChatGPT like it’s the Omnibox in chrome.
I’ve literally seen people ask ChatGPT for a link to Gmail.
Is it a dark pattern, or design decision that makes users happier?
reed1234 17 hours ago [-]
And it hides again and reverts to instant when you are away from the app for a while. Definitely annoying.
Gareth321 13 hours ago [-]
It's getting seriously annoying. This is obviously in an attempt to push people to use the cheaper models but I find them appallingly stupid. There was a worrying comment by Tibo recently where he implied that they're aiming to meter chat queries against one's subscription quota.
krsilas 13 hours ago [-]
I think thats a difference between Chat and Work. In work mode you can select a model and chat should feel "instant".
But I also don't like that differentiation. It's already hard to know which thinking level is necessary, how should users know which mode to use?
maxipoo 17 hours ago [-]
How many can tell the difference?
pants2 17 hours ago [-]
There's a super obvious tell, 5.5 instant loves emojis.
Giving free ChatGPT users access to reasoning (the 'Think' toggle) will have a broader impact on the world than every new paid model and coding agent combined.
daemonologist 1 days ago [-]
Going from 5.5 Instant (which was noticeably bad) to 5.6 Luna is a big jump as well. OpenAI is probably the most prominent among the general public - an advantage in some respects but they're giving away a lot of free inference and thus have to use a pretty small model to do it.
johnsmith1840 1 days ago [-]
I really don't see how they're going to be google here. The free tier is dominated by verticle integration and the platform that people use. Long term I imagine google wins the bottom of the market and I'd be suprised if they lost.
ramraj07 22 hours ago [-]
Even a long while ago it was estimated that every Google search cost google several cents (and that they made back 10 cents). Google also uses AI for search results. Thus its not inconceivable that openai can make money here.
johnsmith1840 21 hours ago [-]
I typed wrong I mean I don't see how they will beat google. Google is the only true vertically integrated AI company meaning they own price at the bottom margins with their own chips and models made just for them.
ekidd 19 hours ago [-]
Google's public AI model lineup is pretty bad right now. Gemini 3.1 Pro is scoring worse than some mid-sized Chinese models at 1/20th the cost, and their Flash and Flash Light models are horrendously overpriced on task benchmarks compared to GPT 5.6 Luna or the mid-sized Chinese models.
There's no reason why Google's public stuff is this stale, overpriced and underwhelming. But at least until their next round of models drops, even calling them a "frontier lab" is starting to feel like a stretch. Which is weird!
willy_k 18 hours ago [-]
That doesn’t matter very much for the free tier market though. The game is to have something useable that is actually used, and that can be offered at sustainable margins. Google has significant advantages in those regards.
eru 18 hours ago [-]
It's interesting that Google fell behind before and managed to get back up to the frontier. But, boy, is their performance in this race uneven.
Especially considering that in the 2010s they were The Big AI company, especially after buying out boutique shops like DeepMind.
tokioyoyo 1 days ago [-]
Ads.
stymaar 1 days ago [-]
Ads work when you can server $.01 worth of ads to a user for $.0001 of server cost. I fail to see how you can make it work when doing LLM inference which is significantly costlier than web search.
famouswaffles 23 hours ago [-]
> I fail to see how you can make it work when doing LLM inference which is significantly costlier than web search.
The median LLM query isn't significantly costlier than web search.
tokioyoyo 1 days ago [-]
In about 4 months, OpenAI’s fundraiser decks will leak, which will include their current ad revenue.
livinglist 20 hours ago [-]
Is this a promise or speculation?
mlvljr 16 hours ago [-]
[dead]
keeda 21 hours ago [-]
That's a problem Google has to solve too ;-)
To expand: it seems inevitable that Google's SERP format will be replaced with a conversational / chatbot / agentic interface, which equalizes the playing field for all chatbot providers.
This is because you can stuff in much fewer ads into a chat interface compared to SERPs. (They could try stuffing more ads but that would likely just push users more to the competition who have a much lower baseline on which to show growth.) As such, Google would be forced to progressively nullify its own invincible firehose of ad revenue as they deprecate SERPs in favor of AI overviews.
red_green_yell 18 hours ago [-]
Google has multiple massive cost advantages over OAI: TPUs, free access to massively valuable data (search index, gmail, maps reviews/PO/ navigation, youtube, android etc.), lower talent costs, lower training costs, lower distribution costs (they can shove AI down our throats in so many different places), lower ad infrastructure costs (it will be a lot of work for OAI to recreate adwords), lower ad sales costs (OAI deploying an ad sales team will be a massive investment).
There's no way OAI has a long term advantage over google in replacing the search engine experience.
Google's one major weakness is it will face the innovator's dilemma as their core search revenue gets cannibalized. But they seem to have been able to get their entire org to recognize that AI is an existential threat so at least that's a good sign.
eru 18 hours ago [-]
I'm not sure Google has lower talent costs. What makes you think so?
red_green_yell 16 hours ago [-]
You are right to question that. I am basing that statement off of the many famous engineers that have recently left and were very highly paid. My guess is Google is ceding the frontier and the high salaries that go with it and letting Anthropic/OpenAI fight over the high priced talent that exit. The remaining non-famous engineers will not be able to command celebrity salaries. But this is just speculation and I don't have great evidence that the celebrity salaries are enough to meaningfully reduce overall talent costs.
A counter point would be OAI and Anthropic can pay with more equity that they can promise will go to the moon. But all the equity base compensation eventually dilutes earnings per share so it's not free once you go public and people start caring about that.
eru 15 hours ago [-]
Agreed. One effect I had in mind was much simpler: if you are the cool new company, people will want to work for you and might even take a hit in salary to do so.
willy_k 18 hours ago [-]
Have you used the Google search engine in the past couple years?
You’re painting a dichotomy that doesn’t exist and hasn’t for a good while.
Edit: I missed your final paragraph. I still disagree with this argument though, people still want to go to websites.
theptip 18 hours ago [-]
OpenAI can very plausibly target ads far better than Google can.
Email is a good window but lots of people talk in more depth with a chat bot.
Search is a good “purchase intent moment” but people are telling a chat bot exactly what they care about in a purchase, rather than making indirect search terms and reading review sites.
aleph_minus_one 9 hours ago [-]
> Email is a good window but lots of people talk in more depth with a chat bot.
> Search is a good “purchase intent moment” but people are telling a chat bot exactly what they care about in a purchase, rather than making indirect search terms and reading review sites.
When I look at my history with AI chatbots, 90 % of the conversations revolve around research-level problems in math and computer science. :-D
theptip 7 hours ago [-]
Yeah I can easily imagine the “AI for work” version never being ad supported and gradually capturing more of the value created. Thats a very different usage pattern.
But 90% (99%?) of user sessions don’t require a huge amount of reasoning tokens or output; keep in mind LLMs are replacing search for single-turn answers.
virgildotcodes 23 hours ago [-]
I recently saw some discussion about the CPC for personal injury lawyers being ~$200 on Google ads. Seems insane to me, but I’m totally disconnected from the advertising world. That said, depending on the conversion rates between chats and clicks, the math could be favorable for OAI.
mikeshi42 22 hours ago [-]
High CACs are supported by high LTVs, that is about the right ballpark for things like lawyers or even plumbers.
kristofferR 24 hours ago [-]
90+% of queries are probably so common/evergreen that, if a cheaper model made all the different language varients and ways of asking the question into one query, they could be cached quite effectively.
aleph_minus_one 9 hours ago [-]
> 90+% of queries are probably so common/evergreen that, if a cheaper model made all the different language vari[a]nts and ways of asking the question into one query, they could be cached quite effectively.
As I wrote in some parallel post:
"When I look at my history with AI chatbots, 90 % of the conversations revolve around research-level problems in math and computer science. :-D"
For quite a few of them, creating a decent answer involved quite a bit of thinking by the AI, so I assume they are not the easiest queries to answer, and the queries are too obscure to cache.
kristofferR 8 hours ago [-]
Sure, but you don't represent the average ChatGPT free user in that case probably. I regularly do super complicated queries that Pro spends ages replying to, but yesterday I also asked if chilis are toxic, or just unpleasant, for dogs, and if other species other than flamingos get colored by the food they eat.
fragmede 22 hours ago [-]
Except that Google is giving AI answers for ~every search now, so they're paying that cost for every search now as well.
theptip 18 hours ago [-]
I bet they cache a lot more than OpenAI can though.
johnsmith1840 21 hours ago [-]
More like LLMs are near commodity at the common tier the ones who wins are those who can inference the cheapest and google is easily thr best positioned to do that with many years of custom chip model optimization.
Google is going to be able to push cost down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.
So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.
That's going to be one hell of a company to beat for free tier LLMs. Also chatgpt is impressive but it's notthing like google search quite yet. On top of that AI labs must use google's product for their AI.
They also own more data by an exponential margin. Anything but dominance of free AI on google's side would mean they are just so incompetent they deserve to fail.
Tostino 1 days ago [-]
People tell these models everything. You don't think a pretty unethical company can figure out a way to monetize that?
aleph_minus_one 9 hours ago [-]
> People tell these models everything. You don't think a pretty unethical company can figure out a way to monetize that?
I assume research-level questions about scientific topics (with a focus on math and computer science) are not easy to monetize (yes, this is by far the most common kind of question to AI models for me). :-D
ewild 8 hours ago [-]
I guarantee you're a lot easier to manipulate than you think.
aleph_minus_one 6 hours ago [-]
> I guarantee you're a lot easier to manipulate than you think.
I can imagine some hypothetical scenarios, but I do believe that I have strong evidence that this is mostly not the case.
For example, I think I have already told the story how when I worked together with a person who is a business consultant (and thus salesman) told me that C-level executives are an incredibly easy sales target compared to me. For him, selling something to me was "ultra-hard mode", basically because I could immediately see through every sales tactic that he tried on me.
fragmede 21 hours ago [-]
If someone tells ChatGPT I'm cheating on my wife, sure, unethically they could blackmail the user but that seems like a stretch.
hluska 21 hours ago [-]
I think the consumer profiling side is more compelling (and likely) than blackmail.
Tostino 20 hours ago [-]
Yup, that is where my mind was.
gavinray 1 days ago [-]
I'm not sure people fully comprehend the trickle-down pop-culture/zeitgeist effects that LLM's are having/are going to have on humanity.
Because everyone now outsources much of their thinking and researching to LLM's, our collective culture + brain is shaped in a cyclical manner by using them.
It's the mechanical homogenization of culture and groupthink.
crab_galaxy 21 hours ago [-]
Is that really accurate though? I use Claude in my software job but the majority of my friends do not use it at all. And I never use it for anything other than that either.
TBH I don’t find it useful at all for personal use. It’s totally soulless for creative ventures and absolute dogshit at researching the things I want it to be good at (I.e. planning a vacation or finding new music).
All that combined with the social stigma makes me feel pretty skeptical that it’s some kind of pop culture shaping mechanism, at least not for a few more years.
fragmede 20 hours ago [-]
We're all in our own personal bubbles and extrapolating from there as normal, but ChatGPT has some 800 million users, most of them outside the US. Countries like the Philippines and Thailand. I don't think they're all vibecoding the hottest new SaaS app but maybe they are.
martheen 17 hours ago [-]
Even the Philippines and Thailand are lower in per-capita usage compared to the US, most of the higher use are in Europe. Interestingly Australia & New Zealand are also pretty high, similar to Europe and far more than Asia & Africa in general, with the exception being Saudi Arabia (similar to US) and Israel (even higher than Australia). https://openai.com/signals/data
14 hours ago [-]
jimbob45 18 hours ago [-]
+1 to this. Many of my engineer friends don’t use AI even in cases where it could clearly help them. Not out of any principles stance but because they just don’t understand that it could speed up their work. It doesn’t help that the advertising surrounding AI hasn’t communicated its benefits well - either you’ve tried it and you get it or you just don’t know. I would imagine some early horseback riders in the 20th century probably felt similarly about the automobile.
chrisweekly 6 hours ago [-]
what industry employs your engineer friends?
in-silico 1 days ago [-]
Does the free tier really not have access to reasoning models?
That would explain a lot of the terrible AI/LLM takes online.
heaney-555 24 hours ago [-]
The free tier of ChatGPT is powered by GPT-5.5 Instant, a non-reasoning version of GPT 5.5.
qingcharles 20 hours ago [-]
Also, remember the bulk of AI users are using free models, getting terrible answers, and wondering why people keep saying AI is going to take everybody's jobs away.
fragmede 20 hours ago [-]
Yes. Go to chat.com in an incognito window and ask it the same question that you ask 5.6-sol with thinking max, and compare. Depending on the question there is or is not a huge difference.
thorum 1 days ago [-]
Free users have had access to reasoning for a while. o4-mini and the initial GPT-5 launch both included reasoning modes for free users.
They took away the button a few months ago and are now putting it back.
heaney-555 1 days ago [-]
Technically yes, but the free usage equated to single-digit messages per day before the auto-downgrade. With GPT-5 it was just 1 message per day.
Perhaps I should have said "proper access".
whazor 14 hours ago [-]
Especially with tool calling.
This could be the Opus 4.5 moment for normal humans.
vitorgrs 17 hours ago [-]
There was already thinking in free model though. They removed the option a while ago and made it automatic... But you could force by just saying "think hard" or whatever.
saithound 1 days ago [-]
While their math results are impressive, vibe coding their own web UIs with their subpar design models is really going to backfire if their plan is to attract new users with better free model offerings.
The Aug 6 update has forced the entry box to auto-format Markdown in an attempt to imitate Claude. The implementation is buggy and even simple copy-and-paste has gone entirely haywire. They also forgot to leave a switch to turn the confounded autoformatting thing off.
Chat mode in general is currently crawling with more UX bugs than a porch screen in summer.
colingauvin 1 days ago [-]
They must really be feeling the commoditization pressure. I'm not sure what the way out of this is, ChatGPT and Claude are still good products, but they are not necessarily premium products anymore.
I expect a few things to happen in the next year:
1) Exclusive MCP server deals/API integrations
2) Significant switch to B2B marketing, even moreso than we've seen before, with API interfaces being paid and chat-client interfaces becoming more and more free, perhaps just with limits more on integrations or data visualization/analysis
3) US restrictions on B2B contracts with non-US hosted models that do any sort of contracting with the government
Obviously there's a bunch of stuff I'm not foreseeing. But it really does feel like the bottom of the market is collapsing into free. I assume OpenAI and Anthropic think their next generation of models will restore their halo tier status and that the cash burn is justified to just get there, but this has to really mess up IPO plans.
redox99 1 days ago [-]
I'm not sure I agree
1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
2) When $100 and $200 per month plans launched, they were received as outrageous even here. Nowadays they are pretty common among power users.
user43928 1 days ago [-]
I also scoffed at the ChatGPT $200 Pro plan back then.
Back when coding for me still meant copy-paste from the web version, it was only worth the $20/month for me.
They only added the $100 Pro plan in April during GPT 5.4 times.
Today I happily pay $400/month for Codex and Claude Code.
thejazzman 23 hours ago [-]
Don’t say the last part out loud or it will be $800.
user43928 16 hours ago [-]
With Kimi K3 and Qwen 3.8 Max on their heels, I think the only way they can pull that off is by selling much better models.
solenoid0937 12 hours ago [-]
I use K3 a lot, it is far behind Fable imo. For personal projects I use Kimi and for anything that matters I use Fable.
colingauvin 1 days ago [-]
>1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
This is kind of what I'm saying though. Bottom has fallen out, differentiation is just can you be much more premium than the competition. Currently that remains unanswered.
EDIT: I'm basing this off the assumption that for chat, premium is not a point of differentiation at all. For coding/analysis, it is.
davidguetta 1 days ago [-]
especially when 99% tasks don't need fable, and certainly not the next model
kingstnap 1 days ago [-]
Its always fun to try to read between the lines here to speculate why they are doing this.
Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
They might be doing this to improve the messaging of AI among causal users since right now there is a huge amount of datacenter backlash in the US due to AI grievances.
Maybe they have too much excess capacity or they really want to juice token numbers and market share on their dashboards for marketing.
I also wonder if being given access to an actually a decent model like luna with actual thinking budget instead of brainless "instant" modes will start to make causal users understand the real capabilities of these models.
planb 1 days ago [-]
Let’s speculate: they want luna to be their first model “on silicon” and need as much test data as possible before finalizing the design.
nialse 10 hours ago [-]
This. It makes sense to develop a quick turnaround silicon process, and it makes sense to do it with the smallest possible "AGI" model which can be served for years in the long tail.
deanc 1 days ago [-]
It's also going to be more training data for them.
ToValueFunfetti 1 days ago [-]
Is this new behavior from them? I haven't looked at their free chat offering in a minute, but I thought they always had them close behind paid tier, often with essentially equal products that made it weird for them to sell the paid tier for chat.
mkozlows 1 days ago [-]
No, free tier was total garbage with GPT-5.
kingstnap 1 days ago [-]
You can't even pick 5.6 luna on the chat app with a paid subscription. It just gives you sol (or use older models) with what seems like basically as much usage as you want. And sol is considerably smarter than luna.
All of this is of pretty minor importance though. You can't read as many tokens as a subcription can produce so more chat is not the value add nor super important.
I mean there are literally so many providers for free chat if you are willing to use several seperate apps.
The real value in these subs is using codex cli, much like the real point of anthropic subs is using claude code. Because agentic work actually does require a lot of tokens.
simianwords 1 days ago [-]
> Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
Definitely this. The recent 80% discount was a reaction to Deepseek's update so that they still position near the frontier. My theory: Luna has always had a much higher efficiency. You do know that the model didn't get faster after the discount?
tosh 1 days ago [-]
free unlimited luna is a pretty badass move
luna is very good
timpera 1 days ago [-]
Any improvement to the ChatGPT free plan is really nice. It's easy to forget that most people have never used a SOTA model, and only think of their experience with GPT-4o or Google Search's AI Overviews when asked about AI.
skybrian 1 days ago [-]
It's also pretty cheap if you're paying for it (for coding). But I don't quite trust it for anything complicated, so I often use Terra.
redox99 1 days ago [-]
Terra is awful and not cheap enough to make up for it. I'd suggest using luna or sol.
stuartq 1 days ago [-]
Terra is far from awful. I've not used Luna enough to judge whether it's significantly better than Luna, but at least via GitHub Copilot, Terra is leaps ahead of Sonnet 5.
drnick1 7 hours ago [-]
It's not awful, but if you are doing research or anything complex and don't want to spend hours fighting it, Sol is not an option.
vitorgrs 17 hours ago [-]
Luna is not bad, but... you feel is a small model. It have less knowledge. It's ok for purely agentic stuff (web searching etc).
Sammi 1 days ago [-]
So Luna is more awful but it's OK because it's cheaper?
LaurensBER 1 days ago [-]
I've had great success with using Luna and having DeepSeek 4 Flash check Luna's output. Oh My Pi has a mode built-in that does this automatically ("advisor" mode). Deepseek only interrupts when it spots an issue so it doesn't slow Luna down.
Both models are cheap enough that I can run 4 sessions at the same time without running out of the 20 USD codex and 10 USD Opencode plan. I've burned through almost a billion tokens this week and I've done some pretty big refactors as well.
I have a Claude Max subscription but I've barely been using it because of the many issues they've had this week.
dannyw 20 hours ago [-]
Yes, in the same way a basic, dry hotdog at $5 is bad, but I could love it in $0.50.
Pricing absolutely matters.
ssl-3 12 hours ago [-]
Indeed.
I buy 2 hot dogs for $1 at Sheetz sometimes. For that price, I get them with mustard, onions, and sauerkraut already on them.
They are not excellent in any way. I will probably never love them. :)
But they're available 24/7/365 and they don't take long for the staff to throw together.
And most importantly: They sure are cheap.
(Costco's $1.50 dog+Coke is much higher quality and presents a better value, but it requires visiting a Costco and that has its own cost.)
redox99 1 days ago [-]
Yes. You're better off running Luna Max or Sol Medium and never touching Terra.
causal 1 days ago [-]
I guess I should give it another shot because I had pretty bad experience with Luna when it first came out. Stuff Sonnet knew better.
beering 4 hours ago [-]
Luna correlates with Claude Haiku, not sonnet. I think Terra would be the equivalent to Sonnet.
msq22 1 days ago [-]
Were free users subject to some limits? I've never run into any limits.
timpera 1 days ago [-]
It used to be 10 messages every 5 hours using GPT-5, then unlimited 4o-mini. More recently, it went down to ~5 messages per day on 5.5 Instant, then unlimited on 5.5 mini.
It’s hard to understand this .. like sure we can select instant but is there an actual model called 5.6 instant? Like is it on LM Arena and OpenRouter or available via API etc
5.5 instant is definitely A Thing it’s even name checked in this OAI post
simonw 1 days ago [-]
It's SO hard to understand this. I couldn't confidently explain it at all.
hluska 20 hours ago [-]
It’s not subtle when you point it out - you’re looking for a different phrase entirely.
1saadcodes 19 hours ago [-]
I dont get why they'd make this available to free users when they're probably dealing with compute constraints considering their competition with the Chinese models. Giving a more capable model to a huge free user base looks like an expensive choice to me
freakynit 19 hours ago [-]
They've probbaly have realized that the foundation models are becoming commoditized (or will be very soon). We can already run luna medium class models on phones now (bonsai, gemma4).
The value is shifting up the stack... make core intelligence free, then monetize the ecosystem built on top of it .. kinda similar to how the internet itself is free, but platforms and apps capture the value.
I think we'll see a huge push towards connectors for work and personal tools, along with much deeper os level integration. That's where the long-term moat is, not the base model itself.
judge2020 18 hours ago [-]
IMO Competition with chinese models frees up higher-turn and bigger-context tasks, since SWEs are more likely to use model routers for coding. Thinking may be a few more steps but the problems are a lot less complex and probably 1/10th the context size, even when skills and memory are at play.
sumedh 11 hours ago [-]
The same reason Google Search is free.
dawnerd 14 hours ago [-]
For the ad integration, obviously. Plus you tell people they can good up their Apple health and now you get more free training data.
ssl-3 12 hours ago [-]
Even though it seems implicit that all of these companies are losing money hand-over-fist, it's still a competitive market.
Improving the free offerings is meant to increase visibility and therefore market share. It's not about goodwill, and it never will be. :)
The free stuff is primarily marketing and marketing always has costs.
In terms of compute, I have no way to really look behind the curtain and see what goes on back there. But I know with codex CLI, in terms of weekly quota: I can get a ton of work done with luna and usually get reasonable results. It feels very compute-light in this way.
I'm amazed by the work luna on xhigh can do for the price I pay (just $20, every month). It has the presentation of something that is very efficient to run, while also being something that can actually produce OK results. It's also fairly quick.
It differs from many previous smaller offerings of yore in this way. Like, I mean: I found stuff like the -mini models and 4o to be utterly useless wastes of my time. Luna isn't like that at all; it can get some stuff done.
So far for me, luna is the most impressive part of the 5.6 rollout. Not because it is best, but because it is useful and cheap.
So if luna is decent (it seems that it is), and if it is in fact light (which seems to be true from what I can observe), and it is offered for free, then it may very well be better, faster, and cheaper than the competition is.
And that's good for visibility. Marketing is all about buying eyeballs.
ElijahLynn 1 days ago [-]
I can't wait to never see a reasoning button ever again. Why do I have to reason about what reasoning level to use?
miki123211 23 hours ago [-]
I have the opposite problem. I'm not well-calibrated on when I'd want lower reasoning than what's available to me (and how to compare that to lower-tier models). OpenAI now has Luna, Terra and Sol, each at Low, Medium, High and Xhigh, with Pro/Ultra depending on harness and plan. That's ~15 possible combinations of model and reasoning level, and there isn't a satisfactory explanation of which one you want for any particular task.
I feel that work is basically split into two tiers, hard (which requires a good model and lots of reasoning by definition) and relatively easy (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High).
majormajor 21 hours ago [-]
> (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High)
This equation significantly changes if you're paying API prices vs on a subscription. (Such as if you're integrating it into a different product, where it becomes worth it to figure out what the cheapest you can leverage is.)
andhuman 16 hours ago [-]
I don’t see why they just don’t allow a smaller model to answer the question while letting the bigger one vet it. The vetting can be asynchronous and can be delivered after a few seconds (if it’s an easy query). If it’s a hard query, the UI can show the answer is currently being vetted or something.
Davidzheng 15 hours ago [-]
if the bigger model can vet it fast it can also answer it fast.
redox99 22 hours ago [-]
Yeah, it's hard, there's so many permutations. But most are bad so it narrows it down.
Luna: Can be useful if price sensitive, always use with AT LEAST high effort. But codex plans are very generous, so just ignore it.
Terra: Forget it exists
Sol: Just use this. Medium can work for specific edits. Otherwise just use high or xhigh.
OpenAI basically agrees with this, and the slider gives you those options.
TL;DR: Just use sol medium/high/xhigh
drivebyhooting 22 hours ago [-]
Why not pro/ultra?
redox99 22 hours ago [-]
Pro: It's web chat only, I haven't tried it so idk. I think it's useful for Math and that kind of stuff, but I only do programming.
Ultra: Might be useful, the one time I tried it, it was extremely wasteful in terms of tokens. Definitely not a daily driver.
There's also max effort, I think it can be useful but the gains compared to xhigh are quite small.
I think max and ultra are only worth it when the others fail.
drivebyhooting 22 hours ago [-]
It’s very confusing to me.
I’ve occasionally used pro to do high-level research and design. Then I ask it to create a prompt for Codex ultra. Ultra can do a lot of genetic benchmarking and testing to elucidate and resolve quandaries.
I don’t know if this is a good workflow.
skybrian 1 days ago [-]
Some people want quick results. Some people want it to keep searching for a new math proof overnight without giving up, and they have money to burn.
It seems like giving it a time limit or a budget in dollars would be clearer, though?
Or, keep searching until I come back to the computer and ask about progress.
pllbnk 1 days ago [-]
The _Intelligence_ part of AGI should be able to guide the user through that without all the knobs.
minimaxir 1 days ago [-]
OpenAI tried auto-routing with the initial GPT-5 release and it was immediately clear why that was a bad idea.
1 days ago [-]
vanuatu 1 days ago [-]
intelligence is not omniscience though
stymaar 1 days ago [-]
Asking questions to clarify user intent is a very low bar for intelligence. A bar that all SOTA models fail consistently at though. (It's both funny and legit infuriating when Opus, after having made a dozen wild assumptions without checking with you, then comes back with a request for clarification on some mundane topic).
vanuatu 22 hours ago [-]
but you can keep asking questions ad infinitum
very nontrivial problem knowing when to stop and making assumptions
stymaar 15 hours ago [-]
Yes, that requires some amount of intelligence, that's exactly my point.
dbbk 24 hours ago [-]
I don't understand this at all. Whenever I ask Gemini 3.1 Pro Extended, or Claude 5 Max something in chat, the most I ever wait is maybe 30 seconds. Is that really so bad?
oceanplexian 23 hours ago [-]
"Wait 30 seconds" as a concept has been totally incompatible with the web, smartphones, etc for about 20 years now.
vitorgrs 17 hours ago [-]
If you just want to know "When it's the next full moon", yes. Very bad as google can answer in 1s.
dbbk 15 hours ago [-]
That I just type into Google.
vitorgrs 3 hours ago [-]
That's also using AI :).
ChatGPT already competes with Google...
nojs 24 hours ago [-]
Auto-effort and similarly auto model routing suffer from a halting problem sort of issue: you don’t reliably know if a request is complex unless you use a complex model to make the decision.
redox99 1 days ago [-]
Because the model can't read your mind and know if you want a quick answer, or an hour long deep dive.
Jtarii 1 days ago [-]
Then it should just ask the user what they want if its unclear from the context.
redox99 1 days ago [-]
So instead of just getting an answer, I have to wait until it asks me, and I have to type back a response? Extremely annoying.
Sammi 1 days ago [-]
That's what the reasoning slider is for!
2sk21 1 days ago [-]
Exactly! I posted much the same comment in another thread and there were lots of huffy complaints that amounted to "you're prompting it wrong"
Marha01 18 hours ago [-]
How is asking better than a reasoning slider?
awakeasleep 1 days ago [-]
Because your incentives are opposed to the provider’s incentives
timpera 1 days ago [-]
I think it's nice to be able to make the model reason for dozens of minutes when you want to go deep on a topic, even if the router thinks it's an easy question.
egorfine 1 days ago [-]
I absolutely need instant mode as this is what I use 90% of the time.
Sadly it's not available anymore on the updated desktop app (formerly Codex) and the previous desktop app (formerly ChatGPT) is abandoned.
visarga 17 hours ago [-]
> I can't wait to never see a reasoning button ever again.
I hope to see a memory-less mode that is not incognito. Want fresh contexts sometimes, but also want to keep the chats saved in history. Memory can spoil some creative work, it dials the model in too tightly.
customguy 22 hours ago [-]
just wait for the loot boxes, mark my words
try-working 16 hours ago [-]
What's even more noticable is that Anthropic still hasn't responded to the Kimi K3 release or the DeepSeek release.
Iolaum 16 hours ago [-]
Do they? With Fable/Opus duo they have a better product and their customers are paying for quality. They want to be the premium LLM provider letting others compete for the commoditized part of the market.
bredren 15 hours ago [-]
Even more so has been what appears to be a miss with Opus 5 in Claude Code.
solenoid0937 12 hours ago [-]
Opus 5 is amazing if you use it as an autonomous agent as opposed to a pair programmer. It's like Fable at a much lower price.
18 hours ago [-]
kaszanka 20 hours ago [-]
> Because this version of GPT‑5.6 Sol is optimized for everyday chats, it will only be available in the Chat experience in ChatGPT. The version of GPT‑5.6 Sol that powers Work and Codex is not changing as part of this release.
Hm, does this mean that 5.6 Pro in ChatGPT web is somehow different/not as good now? I found it really good for code review (upload your repo and patch and off it goes).
itsmeduncan 8 hours ago [-]
How long can OpenAI (and Anthropic) incentivize, and keep costs artificially down for tokens to depress investment in private/offline LLMs? It's interesting to watch the cost curves, and burn to see where it stops being venture subsidizes, and the frontier open-weight models become as good.
qup 7 hours ago [-]
Inference is profitable
Gareth321 13 hours ago [-]
My cynical read on this is that they're going to allow Sol on higher modes to downgrade its mode depending on the question or task.
kgeist 1 days ago [-]
>avoid extra detail when it does not help
I wonder if they actually do it to optimize inference. I maintain a corporate AI server and one of the tricks to reduce the load was to modify the system prompt to be as terse as possible so the average response completes faster and requests queue up less often.
beering 4 hours ago [-]
If you remember how crazy verbose previous gpt models were… clearly there’s something going on unrelated to cost. It would restate the same thing in different words several times and fill the output with emoji or lists.
dannyw 9 hours ago [-]
Oh absolutely! The LLM equivalent of "death and taxes" is prefill and decode; and that holds true irrespective of proprietary inference optimizations.
Less verbose output = less context & less token gen.
lilytweed 15 hours ago [-]
This post makes it sound like paid users should be getting a 5.6 Nonthinking as the Instant option. Still looks like 5.5 Instant to me?
dannyw 9 hours ago [-]
Like most companies and releases, OpenAI doesn't flip the switch for everyone at once. I believe Enterprise seats are also generally on a ~2 week delay from consumer.
Razengan 10 hours ago [-]
> Because this version of GPT‑5.6 Sol is optimized for everyday chats, it will only be available in the Chat experience in ChatGPT. The version of GPT‑5.6 Sol that powers Work and Codex is not changing as part of this release.
So same name, and no indication of which "version" you're using except whether you're in "Chat" or "Work"? Why make it so confusing?
sunaookami 1 days ago [-]
This is actually a downgrade for free users since currently it uses GPT-5.5 for a few messages before it drops you down to GPT-5.5-mini. Now it always uses a model worse than Mini (Luna is nano-equivalent, "It roughly corresponds to the nano model tier used in earlier GPT-5 families." https://developers.openai.com/api/docs/models/gpt-5.6-luna ). I guess it's a bit better with Thinking though. They should use Terra for a few messages first before dropping down to Luna. And image inputs are still limited.
dannyw 9 hours ago [-]
You should try Luna, or look at one of the many independent benchmarks. Even in non-thinking, it's not remotely in the same class as -nano (luna is much smarter).
roytam87 19 hours ago [-]
OpenAI should be more transparent to (free) users about Data Analysis quota and Chat with attachment quota.
OsamaJaber 1 days ago [-]
Expanding free access is mostly an inference cost
Serving cheaply at that scale means routing, batching, and cache hits, not a better model :D
nc55g3g 15 hours ago [-]
free users get unlimited text chats with Luna AND a Think button now? That's a massive upgrade.
applfanboysbgon 1 days ago [-]
> Our mission is to ensure that artificial general intelligence benefits all of humanity. We’re introducing updates to ChatGPT that improve everyday conversations while expanding access for Free users.
My mission is world conquest. I'm writing a comment on an HN thread.
No, those two clauses have no relation whatsoever. I just felt like saying the first sentence because it sounded cool.
solenoid0937 12 hours ago [-]
They are communicating that they have reached AGI in order to create a trail of documentation for the Microsoft legal battle
Squarex 1 days ago [-]
Why pay for ChatGPT Go then?
HDBaseT 22 hours ago [-]
Access to "ChatGPT-Image-2" I suppose?
johnnyApplePRNG 20 hours ago [-]
Excellent.
Yes, please give free users more access and leave paying codex users in the dust.
Wise plan, Sam.
dannyw 9 hours ago [-]
Did you not get all the Codex resets (including banked resets), the removal of the 5hr limit, or even the fact that you can use your Codex sub for API purposes and it's allowed under their terms?
I dunno, I've honestly never been happier as a Codex customer. Sure, there's been less bonus usage resets recently, but the amount of productivity and value I've gotten out of my $20/month personal subscription is bonkers.
If you're price insensitive, Fable 5 is definitely still the best, but a lot of people aren't.
cbg0 15 hours ago [-]
How are they leaving paying users in the dust? This is their cheapest model which you can already use full-blast even on a $20 plan.
beering 4 hours ago [-]
It’s interesting how people will interpret news in ways that seem bizarre, but they were looking for something to be angry about.
nickthegreek 5 hours ago [-]
The literally lowered Luna costs for codex users (and lowered the amount of quota it uses for sub users) earlier this week. This change came for Codex users first.
rrvsh 10 hours ago [-]
What do you mean? Codex paid plan treats me quite well. I'm not going to complain that free users get more shit, especially when it doesn't affect me in any way.
ssl-3 12 hours ago [-]
Eh?
I've paid for ChatGPT (and in the recent ~year, codex) for about as long as it was available to pay for. I'm not upset by the offer of giving luna to the masses for free -- not at all.
Should I be upset that people can cut-and-paste to the lessest of the new model variations for free? If so, then why?
Nothing was taken from me here. I'm still going to keep doing whatever it is that I do with the tools that I've been using.
I'm not in competition with anyone, and even if I were then it wouldn't be with those who are using ChatGPT for free on the web.
ignoramous 1 days ago [-]
Every week, 1 billion people turn to ChatGPT for everything from quick questions and web searches to planning, research, advice, and complex decisions.
Guess, Google's AI Mode is chipping away at their consumers (I know I haven't used Chat in a long, long while for 'quick questions and web searches' after OpenAI did away with "think" which I always use). The money-minting office & coding market Anthropic has cornered is hyper-competitive at both the frontier & low-cost ends. OpenAI is reactive [0] and seems right up against it, despite the strength of its excellent models.
[0] Won't put it past OpenAI (and/or Google) to open weight larger models!
skybrian 1 days ago [-]
$20/month also gives you API access for coding, in any coding agent. I keep hitting the weekly limit but it's a good deal while it lasts.
drivebyhooting 1 days ago [-]
Google’s AI has been very glitchy for me lately. I used to reserve chatGPT for serious work and Gemini for daily personalized unimportant things. But now I switched completely to ChatGPT and resigned myself to their memories/personalization.
sk4rekr0w 19 hours ago [-]
You've jumped to a naive conclusion
porridgeraisin 1 days ago [-]
Did they do away with think? I think now you have to do it with /think
simianwords 1 days ago [-]
Does 5.6 Sol finally have an "instant" form? Is that the change?
nickthegreek 5 hours ago [-]
Instant uses Luna, not Sol.
simianwords 4 hours ago [-]
There’s no luna instant I can see
1 days ago [-]
redox99 1 days ago [-]
Yes
jauntywundrkind 1 days ago [-]
At first I thought the Sol updates was perhaps trying to help with some complaints of Sol burning through tokens, complaints that have prompted some new data points on https://codex-resets.com/ .
But seeing the graphic with the visual weather report: that makes me think that is not the goal at all. :)
laweijfmvo 1 days ago [-]
What a bizarre example of a more direct answer. Any human would simply say “No, it doesn’t rain here in summer.”
Even after identifying the 0% chance of rain, it still drags the conversation on and on and on
taikahessu 1 days ago [-]
Would you like to know more? Seriously though, of course, it's tuned for maximum engagement, not maximum efficiency. I wonder how long this engagement dopamine circus can last... too long apparently.
sunaookami 1 days ago [-]
GPT models were RLHF'd to death, they will never give a final, direct answer. Every release since the GPT-4o catastrophy is like this, it's so tiresome. They need a complete reset before it can become actually usable again. Or not as maximum engagement seems to be their goal.
aniceperson 1 days ago [-]
a m-dash in the title? whoa things are degrading fast
iJohnDoe 18 hours ago [-]
> For Plus and Pro users, we’re updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers.
I was not impressed with 5.6 and this hits exactly why.
Also, this instant, medium, and high slider situation we now have everywhere is batshit crazy. It’s a major step back in technology.
I’ll put money on the table there will be a surprise in revenue because users have no clue what to choose, so they are constantly choosing high because they don’t want to risk getting inaccurate answers. If this was intentional by the dark patterns department, then brilliant. However, I’m guessing Anthropic and OpenAI are struggling to know how to deploy their models.
Also, the models were already amazing. They need to slow down and do a model release once a year and only do extremely minor iterations instead. Some amazing things are accomplished, but how people actually want to utilize the models gets screwed up every time in the process.
LUmBULtERA 5 hours ago [-]
>They need to slow down and do a model release once a year and only do extremely minor iterations instead.
There are a ton of use-cases where the models still struggle a lot and make bad decisions, I'd much prefer they continue their acceleration for a while more.
qhwudbebd 13 hours ago [-]
It seems even worse than that: when you're calling the model, you have this strange two-dimensional thing with reasoning effort (specified in a small handful of random strings like "medium", "xhigh", "max") and "pro" vs "not pro" which is also somehow increasing reasoning budget, but with an interaction that is entirely opaque. How does "medium", "pro" compare with "max", "not-pro" for instance?
If you're going to force people to specify manually, at least make it 0.0 - 1.0 normalised such that 0.5 is the default.
qhwudbebd 13 hours ago [-]
OpenRouter map this into their API in a slightly different way, making -pro and normal different model slugs, so you have openai/gpt-5.6-luna-pro vs openai/gpt-5.6-luna vs openai/gpt-5.6-sol[-pro], etc. This makes reasoning effort one-dimensional again (albeit with the arbitrary sequence of strings), but now model choice in a given generation is two-dimensional. Either way, it's hard to make any kind of informed decision.
julian-vix 10 hours ago [-]
[dead]
porridgeraisin 1 days ago [-]
I dont pay for a chatgpt subscription, but sometimes I did use the web app for throwaway questions. GPT 5.5 Instant or whatever it was that they had was absolutely horrendous. Never answered a question straight and was pedantic in a way even a redditor wouldn't be. So I dropped it and just opened my paid coding agent for everything. Grok.com is quite good now with grok 4.5 though and I find myself using that often. Hopefully luna will be similar.
So 5.6 Luna is just their next version of what they used to call 5.5 instant tier
And 5.x instant models were never much to write home about anyway so the default ChatGPT free model hasn’t been particularly distinctive since 4o
I wonder if this means that Luna is more, uhm, "free tier like" in its responses? The "Instant"/"Chat" models have a pretty particular vibe.
It's also 50% cheaper if you only need it to run sometime in the next 24 hours.
Flex is also easier to get caching to work, there is a little futzing around with OAI's implicit caching but if you do the upfront work you can get haiku quality responses with caching in close to realtime for 1/7th the cost and you don't need to design a polling loop to check for batch completions
This is precisely why I do not offer anything for free. I've been there and done it back in the iPhone 3 era with a dozen apps. Learned my lesson quickly. Never again. It was a nightmare. If the product isn't good enough for someone to pay for it (that could be direct user-paid or advertiser supported) it should not exist.
Also, it is intended to be supported by ads, so it fits your ideal product description quite well despite being a scam.
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.
LLMs still live in the uncanny valley and can be sussed out immediately.
For example, I’m extremely annoyed by the fact that offshore developers respond to me almost exclusively using text generated by Claude. You can tell immediately because they use overly descriptive techno word salad that no normal human uses unless they are trying to be ultra specific for a scientific paper - and even then it’s still too much for a real person.
5 years ago, you would not have been able to determine that this was the case, and just have assumed it's a know it all character.
Tell your offshore developers to use caveman or so
Interesting idea. "This is not dumb and biased enough, probably not a human".
We have one of those: Grok.
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
In the original test the evaluator knew one was a machine and one a human and could have conversations of arbitrary length.
Just look at some training sets to see how the sausage is made: https://huggingface.co/datasets/nickrosh/Evol-Instruct-Code-...
The test wasn't made to accurately measure IQs that high.
They are definitely smart enough to be useful, but dumb enough in their weak spots not to deserve the "general intelligence" qualifier.
As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.
This sounds interesting on it's own. I would be curious to hear more if you are willing to share.
>artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work
[0] https://openai.com/charter/
Take mathematical calculations, styrofoam, LCD screens, ice cubes, embroidery (try searching for "computer work blouses"!), texts, navigation. The thing itself costs nothing, the service and theater around it becomes everything.
Well, the models are smart enough to point out why this is wrong.
Don't you see a problem here? Terms are used to describe the world and need a semblance of stability so we don't end up in a race to the bottom just so investors can feel good.
Do you genuinely hold this position, or do you not realize how far the goalposts have shifted?
In 2022, prominent AI critic Gary Marcus offered to bet $100,000 that we wouldn't have AGI by 2029. https://garymarcus.substack.com/p/dear-elon-musk-here-are-fi... Because the definition of AGI is unclear, he defined that AGI would be achieved if an AI model could do THREE of the five following tasks:
- In 2029, AI will not be able to watch a movie and tell you accurately what is going on (what I called the comprehension challenge in The New Yorker, in 2014). Who are the characters? What are their conflicts and motivations? etc.
- In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI.
- In 2029, AI will not be able to work as a competent cook in an arbitrary kitchen (extending Steve Wozniak’s cup of coffee benchmark).
- In 2029, AI will not be able to reliably construct bug-free code of more than 10,000 lines from natural language specification or by interactions with a non-expert user. [Gluing together code from existing libraries doesn’t count.]
- In 2029, AI will not be able to take arbitrary proofs from the mathematical literature written in natural language and convert them into a symbolic form suitable for symbolic verification.
Today's AI models can do FOUR of these five.
Using 2022 goalposts, we already have AGI. We blew past these goalposts months ago, and nobody noticed.
In the meantime, it’s still very easy to differentiate between an AI and a human in a chat. You just need to know the quirks of these systems. Like counting letters, hitting the safeguards, etc.
So call me when one of them can pass the Turing test against me and then we can talk about AGI
Ring ring I'm calling you right now. We blew past the Turing test goalpost over a year ago, using 2024 models.
https://www.ie.edu/uncover-ie/has-ai-passed-the-turing-test-...
GPT-4.5 passed the Turing Test with a 73% human rating, outscoring actual human subjects. That is, human evaluators considered the AI more human than an actual human, 73% of the time. LLaMa-3.1 was judged to be a human 56% of the time.
The 'strawberry' test was fixed years ago with the invention of CoT; models only fail that test today when thinking is disabled.
On top of that for your strong feelings, you don't have the conviction to write down a strong definition of intelligence yourself, which allows you to accelerate the goal posts up to light speed. The fun thing about writing out a formal definition is suddenly almost everything or almost nothing, including a lot of humans, has intelligence.
Not basing intelligence on your feelings of the moment makes it a hard thing to define across everything intelligence applies to.
Wasn't there a thread the other day about how very few humans on earth can understand the new math proofs?
Although "with sufficient study" vs "not even with unlimited study" are probably worth distinguishing there.
Do you perhaps have a link to it? Tried to search around but couldn’t find it. Would love to read it
OAI is not going to announce anything until their next model officially releases (allegedly later this month).
You can't be serious... Case in point - I just asked GPT Sol High to give me a weekly update of local ai changes.
Here's it's first update, which is complete and utter garbage; i.e. it's a lot of words that says absolutely nothing. That's just a random word generator; AGI? Not even remotely in the ballpark.
https://chatgpt.com/share/6a761ec1-4b38-83ea-a8f7-8f02084e63...
It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?
I’ve literally seen people ask ChatGPT for a link to Gmail.
Is it a dark pattern, or design decision that makes users happier?
There's no reason why Google's public stuff is this stale, overpriced and underwhelming. But at least until their next round of models drops, even calling them a "frontier lab" is starting to feel like a stretch. Which is weird!
Especially considering that in the 2010s they were The Big AI company, especially after buying out boutique shops like DeepMind.
The median LLM query isn't significantly costlier than web search.
To expand: it seems inevitable that Google's SERP format will be replaced with a conversational / chatbot / agentic interface, which equalizes the playing field for all chatbot providers.
This is because you can stuff in much fewer ads into a chat interface compared to SERPs. (They could try stuffing more ads but that would likely just push users more to the competition who have a much lower baseline on which to show growth.) As such, Google would be forced to progressively nullify its own invincible firehose of ad revenue as they deprecate SERPs in favor of AI overviews.
There's no way OAI has a long term advantage over google in replacing the search engine experience.
Google's one major weakness is it will face the innovator's dilemma as their core search revenue gets cannibalized. But they seem to have been able to get their entire org to recognize that AI is an existential threat so at least that's a good sign.
A counter point would be OAI and Anthropic can pay with more equity that they can promise will go to the moon. But all the equity base compensation eventually dilutes earnings per share so it's not free once you go public and people start caring about that.
You’re painting a dichotomy that doesn’t exist and hasn’t for a good while.
Edit: I missed your final paragraph. I still disagree with this argument though, people still want to go to websites.
Email is a good window but lots of people talk in more depth with a chat bot.
Search is a good “purchase intent moment” but people are telling a chat bot exactly what they care about in a purchase, rather than making indirect search terms and reading review sites.
> Search is a good “purchase intent moment” but people are telling a chat bot exactly what they care about in a purchase, rather than making indirect search terms and reading review sites.
When I look at my history with AI chatbots, 90 % of the conversations revolve around research-level problems in math and computer science. :-D
But 90% (99%?) of user sessions don’t require a huge amount of reasoning tokens or output; keep in mind LLMs are replacing search for single-turn answers.
As I wrote in some parallel post:
"When I look at my history with AI chatbots, 90 % of the conversations revolve around research-level problems in math and computer science. :-D"
For quite a few of them, creating a decent answer involved quite a bit of thinking by the AI, so I assume they are not the easiest queries to answer, and the queries are too obscure to cache.
Google is going to be able to push cost down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.
So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.
That's going to be one hell of a company to beat for free tier LLMs. Also chatgpt is impressive but it's notthing like google search quite yet. On top of that AI labs must use google's product for their AI.
They also own more data by an exponential margin. Anything but dominance of free AI on google's side would mean they are just so incompetent they deserve to fail.
I assume research-level questions about scientific topics (with a focus on math and computer science) are not easy to monetize (yes, this is by far the most common kind of question to AI models for me). :-D
I can imagine some hypothetical scenarios, but I do believe that I have strong evidence that this is mostly not the case.
For example, I think I have already told the story how when I worked together with a person who is a business consultant (and thus salesman) told me that C-level executives are an incredibly easy sales target compared to me. For him, selling something to me was "ultra-hard mode", basically because I could immediately see through every sales tactic that he tried on me.
Because everyone now outsources much of their thinking and researching to LLM's, our collective culture + brain is shaped in a cyclical manner by using them.
It's the mechanical homogenization of culture and groupthink.
TBH I don’t find it useful at all for personal use. It’s totally soulless for creative ventures and absolute dogshit at researching the things I want it to be good at (I.e. planning a vacation or finding new music).
All that combined with the social stigma makes me feel pretty skeptical that it’s some kind of pop culture shaping mechanism, at least not for a few more years.
That would explain a lot of the terrible AI/LLM takes online.
They took away the button a few months ago and are now putting it back.
Perhaps I should have said "proper access".
This could be the Opus 4.5 moment for normal humans.
The Aug 6 update has forced the entry box to auto-format Markdown in an attempt to imitate Claude. The implementation is buggy and even simple copy-and-paste has gone entirely haywire. They also forgot to leave a switch to turn the confounded autoformatting thing off.
Chat mode in general is currently crawling with more UX bugs than a porch screen in summer.
I expect a few things to happen in the next year:
1) Exclusive MCP server deals/API integrations
2) Significant switch to B2B marketing, even moreso than we've seen before, with API interfaces being paid and chat-client interfaces becoming more and more free, perhaps just with limits more on integrations or data visualization/analysis
3) US restrictions on B2B contracts with non-US hosted models that do any sort of contracting with the government
Obviously there's a bunch of stuff I'm not foreseeing. But it really does feel like the bottom of the market is collapsing into free. I assume OpenAI and Anthropic think their next generation of models will restore their halo tier status and that the cash burn is justified to just get there, but this has to really mess up IPO plans.
1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
2) When $100 and $200 per month plans launched, they were received as outrageous even here. Nowadays they are pretty common among power users.
Back when coding for me still meant copy-paste from the web version, it was only worth the $20/month for me.
They only added the $100 Pro plan in April during GPT 5.4 times.
Today I happily pay $400/month for Codex and Claude Code.
This is kind of what I'm saying though. Bottom has fallen out, differentiation is just can you be much more premium than the competition. Currently that remains unanswered.
EDIT: I'm basing this off the assumption that for chat, premium is not a point of differentiation at all. For coding/analysis, it is.
Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
They might be doing this to improve the messaging of AI among causal users since right now there is a huge amount of datacenter backlash in the US due to AI grievances.
Maybe they have too much excess capacity or they really want to juice token numbers and market share on their dashboards for marketing.
I also wonder if being given access to an actually a decent model like luna with actual thinking budget instead of brainless "instant" modes will start to make causal users understand the real capabilities of these models.
All of this is of pretty minor importance though. You can't read as many tokens as a subcription can produce so more chat is not the value add nor super important.
I mean there are literally so many providers for free chat if you are willing to use several seperate apps.
The real value in these subs is using codex cli, much like the real point of anthropic subs is using claude code. Because agentic work actually does require a lot of tokens.
Definitely this. The recent 80% discount was a reaction to Deepseek's update so that they still position near the frontier. My theory: Luna has always had a much higher efficiency. You do know that the model didn't get faster after the discount?
luna is very good
Both models are cheap enough that I can run 4 sessions at the same time without running out of the 20 USD codex and 10 USD Opencode plan. I've burned through almost a billion tokens this week and I've done some pretty big refactors as well.
I have a Claude Max subscription but I've barely been using it because of the many issues they've had this week.
Pricing absolutely matters.
I buy 2 hot dogs for $1 at Sheetz sometimes. For that price, I get them with mustard, onions, and sauerkraut already on them.
They are not excellent in any way. I will probably never love them. :)
But they're available 24/7/365 and they don't take long for the staff to throw together.
And most importantly: They sure are cheap.
(Costco's $1.50 dog+Coke is much higher quality and presents a better value, but it requires visiting a Costco and that has its own cost.)
(This is a subtle nudge at anyone from OpenAI who reads this to make sure they get updated.)
OpenAI have a model called "chat-latest" - I wonder if that's running this new model yet: https://developers.openai.com/api/docs/models/chat-latest
It's described as "points to the latest Instant model currently used in ChatGPT" - so presumably that's "GPT-5.6 Instant" in the app.
https://gist.github.com/simonw/aae4febd3c6f7bc5b7811857edb3c... has screenshots that still show "Instant" as an option for ChatGPT Chat... but not for ChatGPT Work.
It’s hard to understand this .. like sure we can select instant but is there an actual model called 5.6 instant? Like is it on LM Arena and OpenRouter or available via API etc
5.5 instant is definitely A Thing it’s even name checked in this OAI post
The value is shifting up the stack... make core intelligence free, then monetize the ecosystem built on top of it .. kinda similar to how the internet itself is free, but platforms and apps capture the value.
I think we'll see a huge push towards connectors for work and personal tools, along with much deeper os level integration. That's where the long-term moat is, not the base model itself.
Improving the free offerings is meant to increase visibility and therefore market share. It's not about goodwill, and it never will be. :)
The free stuff is primarily marketing and marketing always has costs.
In terms of compute, I have no way to really look behind the curtain and see what goes on back there. But I know with codex CLI, in terms of weekly quota: I can get a ton of work done with luna and usually get reasonable results. It feels very compute-light in this way.
I'm amazed by the work luna on xhigh can do for the price I pay (just $20, every month). It has the presentation of something that is very efficient to run, while also being something that can actually produce OK results. It's also fairly quick.
It differs from many previous smaller offerings of yore in this way. Like, I mean: I found stuff like the -mini models and 4o to be utterly useless wastes of my time. Luna isn't like that at all; it can get some stuff done.
So far for me, luna is the most impressive part of the 5.6 rollout. Not because it is best, but because it is useful and cheap.
So if luna is decent (it seems that it is), and if it is in fact light (which seems to be true from what I can observe), and it is offered for free, then it may very well be better, faster, and cheaper than the competition is.
And that's good for visibility. Marketing is all about buying eyeballs.
I feel that work is basically split into two tiers, hard (which requires a good model and lots of reasoning by definition) and relatively easy (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High).
This equation significantly changes if you're paying API prices vs on a subscription. (Such as if you're integrating it into a different product, where it becomes worth it to figure out what the cheapest you can leverage is.)
Luna: Can be useful if price sensitive, always use with AT LEAST high effort. But codex plans are very generous, so just ignore it.
Terra: Forget it exists
Sol: Just use this. Medium can work for specific edits. Otherwise just use high or xhigh.
OpenAI basically agrees with this, and the slider gives you those options.
TL;DR: Just use sol medium/high/xhigh
Ultra: Might be useful, the one time I tried it, it was extremely wasteful in terms of tokens. Definitely not a daily driver.
There's also max effort, I think it can be useful but the gains compared to xhigh are quite small.
I think max and ultra are only worth it when the others fail.
I don’t know if this is a good workflow.
It seems like giving it a time limit or a budget in dollars would be clearer, though?
Or, keep searching until I come back to the computer and ask about progress.
very nontrivial problem knowing when to stop and making assumptions
ChatGPT already competes with Google...
Sadly it's not available anymore on the updated desktop app (formerly Codex) and the previous desktop app (formerly ChatGPT) is abandoned.
I hope to see a memory-less mode that is not incognito. Want fresh contexts sometimes, but also want to keep the chats saved in history. Memory can spoil some creative work, it dials the model in too tightly.
Hm, does this mean that 5.6 Pro in ChatGPT web is somehow different/not as good now? I found it really good for code review (upload your repo and patch and off it goes).
I wonder if they actually do it to optimize inference. I maintain a corporate AI server and one of the tricks to reduce the load was to modify the system prompt to be as terse as possible so the average response completes faster and requests queue up less often.
Less verbose output = less context & less token gen.
So same name, and no indication of which "version" you're using except whether you're in "Chat" or "Work"? Why make it so confusing?
Serving cheaply at that scale means routing, batching, and cache hits, not a better model :D
My mission is world conquest. I'm writing a comment on an HN thread.
No, those two clauses have no relation whatsoever. I just felt like saying the first sentence because it sounded cool.
Yes, please give free users more access and leave paying codex users in the dust.
Wise plan, Sam.
I dunno, I've honestly never been happier as a Codex customer. Sure, there's been less bonus usage resets recently, but the amount of productivity and value I've gotten out of my $20/month personal subscription is bonkers.
If you're price insensitive, Fable 5 is definitely still the best, but a lot of people aren't.
I've paid for ChatGPT (and in the recent ~year, codex) for about as long as it was available to pay for. I'm not upset by the offer of giving luna to the masses for free -- not at all.
Should I be upset that people can cut-and-paste to the lessest of the new model variations for free? If so, then why?
Nothing was taken from me here. I'm still going to keep doing whatever it is that I do with the tools that I've been using.
I'm not in competition with anyone, and even if I were then it wouldn't be with those who are using ChatGPT for free on the web.
[0] Won't put it past OpenAI (and/or Google) to open weight larger models!
But seeing the graphic with the visual weather report: that makes me think that is not the goal at all. :)
Even after identifying the 0% chance of rain, it still drags the conversation on and on and on
I was not impressed with 5.6 and this hits exactly why.
Also, this instant, medium, and high slider situation we now have everywhere is batshit crazy. It’s a major step back in technology.
I’ll put money on the table there will be a surprise in revenue because users have no clue what to choose, so they are constantly choosing high because they don’t want to risk getting inaccurate answers. If this was intentional by the dark patterns department, then brilliant. However, I’m guessing Anthropic and OpenAI are struggling to know how to deploy their models.
Also, the models were already amazing. They need to slow down and do a model release once a year and only do extremely minor iterations instead. Some amazing things are accomplished, but how people actually want to utilize the models gets screwed up every time in the process.
There are a ton of use-cases where the models still struggle a lot and make bad decisions, I'd much prefer they continue their acceleration for a while more.
If you're going to force people to specify manually, at least make it 0.0 - 1.0 normalised such that 0.5 is the default.