I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Btw, they had 3D pelican-on-the-bicycle easter egg in one of the promo videos: https://youtu.be/bOC3DisEOfg?t=117 so I'm pretty sure that they spent some small amount of resources to train the model to produce good svg version as well. :D
theturtletalks 14 hours ago [-]
This is one of those benchmarks I don’t mind if they benchmax because it would mean models are actually good at creating svg graphics.
theturtletalks 13 hours ago [-]
Look at how well it made this Xbox controller SVG:
Do you know what reasoning level this was generated at?
SkiFire13 2 hours ago [-]
No, it would mean models are good at pelicans svgs, not svgs in general.
alexgoodhart 5 hours ago [-]
I’m not sure that this follows. And I don’t understand why the pelican bro is not purposefully demonstrating variety of svg tasks to begin with
kyorochan 9 hours ago [-]
The 3D version is interesting, because it's more detailed which means more things to get wrong. The mudguards are symmetrical for some reason which you would never see in real life, and there are three brake cables but no brakes!
I also think there's an extra level that I would hope an AI would nail which maybe an amateur artist would also fail at, such as thinking about what position a pelican would actually ride a bike in (maybe angling the beak down for aerodynamics etc.), but we are far away from this.
stymaar 15 hours ago [-]
When you see how good the output of the Luna model without reasoning is compared to SotA just a year and a half ago, it's pretty clear that it's been trained on explicitly.
maleldil 14 hours ago [-]
I don't see how this is proof, and not just that the model got the better. You're comparing models a year and a half apart; this is a lifetime in LLM development.
stymaar 8 hours ago [-]
It's a lifetime on things that are explicitly being trained on! But small models like Luna didn't magically become more powerful than SotA models on stuff that they weren't explicitly trained on with a dedicated RL-pipeline.
brookst 12 hours ago [-]
Yeah, suggestive but hardly proof
WarmWash 10 hours ago [-]
Many people still believe that LLMs can only reproduce what is in their training set.
stymaar 8 hours ago [-]
That'd be too restrictive, but to get improvement in a particular domain you definitely need to train specifically for it. Just cramming more random internet text in a bigger model has stopped being a effective way of scaling since at least mid 2023.
threatripper 21 hours ago [-]
It seems to start understanding how bicycles work. Look at the front fork. There are reasons why it's shaped the way it is in real bicycles. If you do physical simulations in 3D with reinforcement learning you will start understanding that most of the arrangements that the other agents use just won't work.
epihelix 20 hours ago [-]
Yes, the curve of the front fork on the max pelican was what I noticed also. It's impressive (if not an accident), and a remarkably accurate bike overall. The pelican just needs to raise their seat a little.
jayd16 5 hours ago [-]
Probably just luck seeing at the wheel intersects the frame and the penguin's legs are caught in the chain.
steve-atx-7600 23 hours ago [-]
You would not expect the developers of the model to optimize for a well known benchmark?
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
swingboy 41 minutes ago [-]
It seems like Astra has a color palette it likes.
Fuzzwah 4 hours ago [-]
It did that thinking about a helmet but didn't actually include a helmet. And there's no "actually, wait, this is just a cute little image no helmet" thought.
pilaf 18 hours ago [-]
Interesting how the background is almost identical to two of the Astra pelicans'.
stymaar 14 hours ago [-]
What if you asked a front/rear/top-side view of the scene (pelican or lemur)?
Did even better from the front. What's surprising is that it used the exact same colors as the simonw example, despite my prompt only being
> Generate an SVG of a ring-tailed lemur riding an electric scooter. Front view
GPT-6 Astra Extra High
15 hours ago [-]
benatkin 19 hours ago [-]
Hmm, the face makes it not work as a one shot artifact.
CamperBob2 20 hours ago [-]
If you want to see something brutal, ask for a zebra riding a scooter. I have yet to see any models do a credible job at that.
ulrikrasmussen 16 hours ago [-]
I remember that early image models couldn't generate a cyclops no matter how you prompted it, it would at best put a third eye in the forehead.
pointitkememe 17 hours ago [-]
[flagged]
boxed 14 hours ago [-]
The lemur is worse though. The hands don't follow physics for example. You claim this would be easier, but then the model is worse.
y1n0 23 hours ago [-]
I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.
Kranar 23 hours ago [-]
What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?
mudkipdev 22 hours ago [-]
Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.
awakeasleep 22 hours ago [-]
You gotta look up how RLHF works before you ask a demanding question like this.
WarmWash 10 hours ago [-]
Right, but that's not the crazy part. The crazy part is thinking they do it all specifically for pelicans on bikes.
If that was the case, the models would have been producing near perfect outputs for it a year ago.
Instead they are just training on general SVG generation, which in no way should be viewed as "benchmaxxing".
dgellow 17 hours ago [-]
> Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles?
Yes
jenniferhooley 11 hours ago [-]
Do you really think there aren't data annotator services specifically training for SVG drawing? And that those annotators don't have a rich set of frequently requested icons/graphics/etc. that they review and train on?
I don't think it's a great rebuttal, quite the opposite actually: it's the same prompt with the most obvioy tweak you can imagine and the results are all the same. That definitely looks like it's been something the models have been trained on, and not some emergent capability (and if you take into account how nonthinking Luna compares to the SotA models from a year an half ago: https://simonwillison.net/tags/pelican-riding-a-bicycle/?pag... it becomes even more obvious)
boxed 14 hours ago [-]
How about you come up with something yourself instead of just moving the goalpost?
stymaar 8 hours ago [-]
No goalpost was harmed in the above comment.
mi_lk 22 hours ago [-]
It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it
Treat it like a bit as is
benatkin 21 hours ago [-]
It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.
epiccoleman 12 hours ago [-]
Starting my SVG of a pelican on a bicycle as a service startup today, invest now for infinite returns!
justinbaker84 15 hours ago [-]
Your comparison grid is one of the most useful comparison grids I have ever seen in terms of measuring AI.
It feels silly to say that about making a pelican image but it really shows the difference in output and cost in an easy to understand way.
jsdalton 1 days ago [-]
Simon do you have a page somewhere that shows _all_ of the penguins created by various models across time?
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
jostylr 6 hours ago [-]
You could also have it create a viewer competition selection; show two images without attribution and let the use select which one. Could be cool to see the wisdom of the crowds; could also have price ranges to have subcategories. I bet Astra could do that in a half hour.
Loving the pelican silliness. My ChatGPT is over the moon about it. Gonna miss it when it really is done. Though maybe a pelican riding a bike game could become the benchmark in a year.
I used Light mode and it used up all my limits for the day and had to continue the following day.
CamperBob2 6 hours ago [-]
This is great! I can see some of the more abstract ones ending up on the wall at SF-MoMA.
Missing GLM-5.3, though.
bahmboo 21 hours ago [-]
Would be great to spend some money (other people's money) on paying some illustrators and artists to do the same for comparison. It could be straight images converted to SVG or direct vector graphics.
WarmWash 10 hours ago [-]
Technically you would pay programmers to write SVG code with no visual feedback if you want a straight 1:1 comparison.
bahmboo 5 hours ago [-]
Kinda. It's the difference between measuring the end product vs how it's made.
Gander5739 18 hours ago [-]
Doesn't that somewhat miss the point? The models can't see and check their work, for instance.
bahmboo 17 hours ago [-]
I'm not clear what you are saying. I am curious about what humans can do in the same context. It's becoming a standard I think for some tasks. E.g. how does a Gen AI stack up against a human in quality and "cost". Perhaps I'm missing what you are asking.
spockz 14 hours ago [-]
They can. Ask them to generate SVG , call a renderer and let them inspect the bitmap/png. They can learn about it.
Gander5739 11 hours ago [-]
They can, but for this specifc test, they don't. It's a oneshot with no feedback.
vb-8448 1 days ago [-]
I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.
thimabi 1 days ago [-]
This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.
pizza234 1 days ago [-]
Astra is the only model that correctly depicts occlusion of crank and leg, although interestingly, at max and medium levels (not in between).
vessenes 23 hours ago [-]
Looked to me like it missed the chain though.
TiredOfLife 13 hours ago [-]
Kimi k3 and Fable max also
petilon 23 hours ago [-]
Astra pelican looks amazing: it looks like it was made by a professional artist. All the others look like they were made by kindergartners.
andai 1 days ago [-]
Oh my goodness. I was not prepared for Luna on "none".
It's so meme-worthy next to the astra-max for any situation of "What you were promised" vs "What you received".
tyre 1 days ago [-]
Haiku the GOAT
jeffybefffy519 18 hours ago [-]
Shouldnt you give the AI a different task each time? otherwise the model companies just optimise for this benchmark because its in their best interest to...
mkagenius 19 hours ago [-]
Do you ever randomize on a particular model - like try and get 3 outputs and pick one at random?
Coz who knows if astra low will produce max like output if tried once more.
BrokenCogs 1 days ago [-]
Interesting that all of the bikes are turquoise colored, except for the medium effort
alastairr 19 hours ago [-]
It's mildly interesting that the bikes always seem to be front wheel facing the right.
simonw 11 hours ago [-]
Most images of bicycles on the web are arranged that way because the drivetrain is always on the right of the bicycle and you want that displayed prominently.
alastairr 11 hours ago [-]
I don't know how I didn't think of that! Makes sense.
1 days ago [-]
mkl 11 hours ago [-]
Why do Sol and Terra use 26 input tokens when the others use 16?
simonw 10 hours ago [-]
I don't know and I'm really
Interested in the answer.
I heard a rumor that Luna is a slightly different architecture from Sol and Terra, which makes me wonder if Luna and Astra might be more related to each other than to Sol and Terra.
Bit of a big leap to make from a token count though!
The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.
scotty79 14 hours ago [-]
They really can't decide which way pelican's knees should bend, do they?
EugeneOZ 15 hours ago [-]
I love Astra pelicans! Awesome results!
fHr 24 hours ago [-]
Luna is my go to daily model it's great value
tuo-lei 23 hours ago [-]
Me toooo, for all my personal projects I need to pay for the tokens~
At workplace I use sol because I don't need to pay
ghthor 24 hours ago [-]
Mine as well, it’s fast and keeps me in flow; and cheap!
poopiokaka 19 hours ago [-]
[dead]
man4 23 hours ago [-]
[dead]
jedjjfjf 22 hours ago [-]
[dead]
leoqa 1 days ago [-]
[flagged]
samuelknight 1 days ago [-]
How are we supposed to know if Astra is frontier without the pelican?
satvikpendem 1 days ago [-]
It's simonw. It's interesting to see their pelican benchmark, another comment by a different author elsewhere here shows some very good SVG generation too.
ComplexSystems 1 days ago [-]
It's absolutely related.
StopTheCringe 1 days ago [-]
[flagged]
bnorton 23 hours ago [-]
You’re closer to this than I am but do you find it odd that helmets are almost never included? Biking is almost always accompanied by helmets
maxlapdev 23 hours ago [-]
But pelicans are almost never accompanied by helmets, so it cancels out.
sumedh 23 hours ago [-]
> Biking is almost always accompanied by helmets
Isnt Netherlands the leader in bike riders and they dont wear helmets.
neutronicus 23 hours ago [-]
At least one reasoning trace I saw considered a helmet and discarded the idea because it was worried about obscuring some detail
frenchtoast8 10 hours ago [-]
I signed up for OpenRouter, loaded it up with $25, and after running a test prompt immediately had my account suspended. There’s no way to talk to support, emails go nowhere, and their Discord is swarmed with people who also are getting no responses to anything.
I would recommend staying away from OpenRouter. No matter how good the service is, if anything does go wrong, you have no recourse and you lose every credit in your account. Ironically some of the few responses I actually saw in the Discord were doubling down on their “no refunds no matter what” policy.
Gecko4072 8 hours ago [-]
They were just bought for $7B and rely on Discord?
bellowsgulch 7 hours ago [-]
God, that's a step up from AI support, which is so sad to type out loud.
miyuru 10 hours ago [-]
take screenshots of the responses and do a chargeback.
predkambrij 9 hours ago [-]
I did get a response when sending to support@openrouter.ai about a year ago (complaining about some information claims), not about my account.
frenchtoast8 6 hours ago [-]
I emailed them a few days ago and the automated reply says to expect a week for a response. That wouldn’t be concerning except on Discord there are people begging for help after waiting multiple weeks.
tocariimaa 8 hours ago [-]
Are there alternatives that don't slurp in the remaining unused credits every new month? OpenRouter credits carry to the next month up to a year so that's why I use it.
weberer 6 hours ago [-]
AWS Bedrock is good with billing and customer service, but they don't offer as many models.
8 hours ago [-]
dominick-cc 8 hours ago [-]
That's crazy. Thanks for the heads up.
8 hours ago [-]
enraged_camel 8 hours ago [-]
Yeah their support is non-existent. It's actually mind-blowing that so many people use it.
bellowsgulch 7 hours ago [-]
Thanks, might move our corporate accounts over to proxy billing now.
jjcm 22 hours ago [-]
It's ability to handle non-90 degree cutouts and shapes for web dev is one of the best I've seen. The vision model on this is VERY capable.
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
copperx 19 hours ago [-]
> That site build cost $24 - extremely non-trivial for a simple frontend.
I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.
jermaustin1 8 hours ago [-]
A bespoke design like that a year ago would have cost $200-2000+ between design and development. Hell, it probably still does unless the person wanting the design is already a developer who knows how to prompt.
Everything costs more now than it did a year ago... except for THIS, and we are still complaining that a 90-99% reduction in cost is STILL too expensive. And a 50-75% reduction in time is STILL too long.
We used to have to wait for weeks for a design like that when I worked at a consultancy, and that is a week of salary. For the design, then it got handed off to a front end developer to slice it and get built so the back end developer can hook it up to a CRM. We are talking a month turn around with design, revisions, development, testing, and bug fixing.
It can now be done in a couple of hours for less than a single hour's cost. If it were 10x slower and 10x more expensive, it would STILL be "good deal".
kolinko 14 hours ago [-]
Wdym no room for error or experimation? Changes are even cheaper, and in my experience they are faster and way cheaper with models than with designers.
paxys 9 hours ago [-]
Plus you can offload a lot of work to a cheaper model
JumpCrisscross 16 hours ago [-]
> truth is that the cost doesn't leave much room for error or experimentation
Compared to what?
copperx 6 hours ago [-]
A cheap Chinese model, of course.
risyachka 10 hours ago [-]
>> The truth is that the cost doesn't leave much room for error or experimentation
Yeah if you are solo developer without budget.
For any business this is nothing, the ROI is massive.
mydreamof 18 hours ago [-]
I don't get it. For me it seems Opus was more accurate in terms of for example this small building in the right down corner
CapsAdmin 17 hours ago [-]
I would say opus was in some ways more accurate, but missed the higher level curvature feel of the site that astra picked up on.
It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature
esikich 18 hours ago [-]
They look pretty similar to me, I'm not sure what you're seeing that is so much better.
w4yai 18 hours ago [-]
Also the curves and background lines that separate the right image from the left content
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
satvikpendem 1 days ago [-]
That's quite shocking, at a sufficiently advanced level we can make all non-realistic graphics purely out of SVGs, as they'd have good scaling for things like logos and app icons. I know it was technically and theoretically possible before AI but most people weren't spending hours tweaking SVG HTML. I remember making an SVG dark mode toggle icon and it took days to get it right, I assume it's one shottable now.
dprkh 23 hours ago [-]
I thought all the designers have been using vector graphics for a long time now.
embedding-shape 15 hours ago [-]
"Designer" is a very broad term and can mean very different things to different people, not everyone is doing stuff that can be represented nicely by vectors.
dprkh 4 hours ago [-]
What's the term for the kind of designer who does vector graphics? Graphic designer?
mceachen 22 hours ago [-]
Aldus Freehand 1.0 was released in 1988. Adobe Illustrator was first released in 1987.
kulahan 1 days ago [-]
That photo looks pretty hilariously stupid, so this appears to be more of a first toe dip rather than some indication we can one-shot a previously difficult process.
appplication 1 days ago [-]
Sure, it is a bit cartoonish but it’s relatively impressive. I do wonder what you would get if you asked for photorealism
redox99 21 hours ago [-]
The problem is not photorealism. The SVG is outright dumb, its on the wrong side of the table, the table has fucked up geometry (its tilted) and many more minor flaws.
XCSme 17 hours ago [-]
That's true, if it really "knows" stuff, a bit weird to have basic flaws like 2 mouths, and sitting on the side...
embedding-shape 1 days ago [-]
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to indicate.
XCSme 1 days ago [-]
Do you prefer the fable one?
It's more "correct" but looks a lot worse in my opinion:
And not to mention, the hamster is standing at the wrong end of the table.
sehugg 15 hours ago [-]
And also the hamster has two mouths
embedding-shape 12 hours ago [-]
I think it might be standing at the wrong edge of the table too, worth adding.
embedding-shape 15 hours ago [-]
Incredible, I've played table-tennis since I was like 14, and I didn't notice that egregious error yet I noticed all the other small ones! Thanks for pointing that out, should have been very obvious.
XCSme 1 days ago [-]
Yeah, that's confusing, the score is for the entire benchmark, not for SVG generation only.
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
I've replaced "Score" there with model ranking, to reduce confusion, thanks for the feedback!
embedding-shape 15 hours ago [-]
Now I'm wondering why it's ranked #4 instead, not sure this reduced confusion :P
What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all
XCSme 15 hours ago [-]
Yeah, ranking is for entire model, not only SVG generation. Maybe I remove it entirely and add generation date instead.
I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.
I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).
readams 1 days ago [-]
The Astra one looks pretty good except it's standing on the wrong side of the table
viccis 19 hours ago [-]
>AGI
>Playing ping pong from the side of the table
I think marketing might be getting a bit absurd at this point
mawadev 16 hours ago [-]
I must live in a different world or it feels like I'm reading LLM generated comments, but who exactly is falling for this marketing?
I even keep seeing obvious stealth marketing like this: "<topic> and how do I use it with <product> in <product>"
jubilanti 10 hours ago [-]
> who exactly is falling for this marketing?
The entire mainstream media and political establishment, and every normie I meet is convinced Skynet is already upon us
ryanschaefer 8 hours ago [-]
For me it’s the two mouths
1 days ago [-]
13 hours ago [-]
MisterMunchkin 19 hours ago [-]
$10/$50 is incredibly expensive compared to Chinese models which are cents.
I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
kolinko 13 hours ago [-]
It's like saying you're not paying for senior level engineers because you can get juniors for quarter the price per hour.
wilkystyle 6 hours ago [-]
Assuming what GP said is accurate (tokenmaxxers not delivering more value for the usage), the analogy is more like realizing the seniors you hired aren't actually producing any better results than juniors you could get for a quarter of the cost.
bigyabai 7 hours ago [-]
If the assignment isn't a 600,000 SLOC greenfield superproject, that's a perfectly acceptable alternative.
sschueller 12 hours ago [-]
Not like companies to fire senior engineers and then hiring two juniors to replace them to optimize expenses... /s
kolinko 11 hours ago [-]
An official motto of a president of one of the biggest software companies in Poland (Comarch) was "you can replace any experienced engineer with a finite amount of sutdents... for half of the cost".
It must be super interesting working there rn :D
wookmaster 8 hours ago [-]
My company literally mandated tokenmaxxing while all the engineers told them this was a bad idea.
gentlewater 17 hours ago [-]
Not really comparable IMO. Astra and Fable are not the every day workhorse you reach for to do basic tasks (unless your company has fuck you-money), they’re the tool you break out when you need the absolute strongest performance. There are plenty of tasks where finding and fixing one or two extra edge cases saves the business a lot of money, even if the cost is high. The best example would be scanning for vulnerabilities, if these models weren’t kneecapped in that area.
ghosty141 16 hours ago [-]
We have ChatGPT Pro at work and I usually use Terra medium/high and only bring out Sol High when the big or feature actually requires "thinking"/complex behavior. This has worked pretty well for me and it's very token efficient
yurishimo 7 hours ago [-]
How much code are you shipping in a day? I find I can pretty comfortable use Sol high most of the day and stay within the 5 hour limit. I’ve got too many meetings to allow me time to write code continuously for an entire day. I usually finish about one ticket a day and then review 1-3 tickets for my colleagues.
ghosty141 7 hours ago [-]
We have a very lean development process and are in early stages of releasing the software, so almost no meetings and 90% of my time is development time (obv I talk with colleagues etc in that time). For example working with yocto eats through tokens like crazy since codex has to work with a pretty big codebase and look up tons of stuff.
KptMarchewa 18 hours ago [-]
the only thing that matters is cost per task. Astra seems to be massively efficient.
simianwords 18 hours ago [-]
No it doesn't. Any source?
saidnooneever 12 hours ago [-]
in the middel of a coding session with 5.6 sol. astra popped up, swapped, asked review, it fixed a few really critical bugs immediately. pretty nice. one around some resource lifetimes in a rendering pipeline which would have been a nightmare to find manually.
almost had the feelin it was watching its little brother fail and had to 'step in' for a moment :').
time to go play outside...
kzrdude 7 hours ago [-]
The names, I thought they were going to keep Sol, Terra, Luna for a while (while increasing versions). Are names like Sol and Astra really burned as one-offs? I think that's a waste of a good model name. Hope they keep them going with updates, as nicknames for the various model sizes on offer.
kdnvk 6 hours ago [-]
Astra is a fourth tier above Sol, analogous to Fable and Opus.
kingstnap 1 days ago [-]
Its also available finally to Pro users! Just took 24 hours.
InsideOutSanta 1 days ago [-]
They gave out bankable resets for every day people on pro plans didn't get Astra. Given that, I wish they'd waited a few more days before activating it on my account :-D
embedding-shape 1 days ago [-]
> They gave out bankable resets for every day people on pro plans didn't get Astra.
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
wincy 10 hours ago [-]
I must have been one of the last ones to get it because I managed to snag two bankable resets. Already reset once after having it clear up geometry in Blender, but it did a fantastic job!
wincy 1 days ago [-]
They haven’t activated Astra for me yet, I have two resets now. I’ve been using the opportunity to test out how good 5.6 Sol is at computer use asking it to generate stuff in Blender which has been… interesting
Edit: nevermind it JUST gave me a notification to use it!
paxys 1 days ago [-]
That's a pretty genius internal incentive to move fast.
killerstorm 11 hours ago [-]
I tried it with some humanities questions and with the default OR system prompt (no prompt?) it seems to be rather mild - lacking usual AI mannerisms. Kinda cool.
cmrdporcupine 5 hours ago [-]
I'm appreciating Astra writes much more humanely somehow than Sol did, and certainly better than the Anthropic models.
It's terse, like all GPT models by default, but the sentences feel less obscurantist.
It's also more pro-active about problem solving.
theagenticleade 5 hours ago [-]
This feels like a genuine step change in AI development. Inline with Dec 2025 release of Opus 4.6.
Where do we go from here??
sejje 4 hours ago [-]
Faster; cheaper.
sumedh 23 hours ago [-]
Just got access to it on Plus plan in Australia. 2 Banked resets as well.
1saadcodes 17 hours ago [-]
The higher price seems less important if it actually gets the job done with fewer tokens. I'm still very worried that this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all
vb-8448 1 days ago [-]
Played in codex app a couple of hours today: it feels much faster than SOL, even if the TPS is half of it.
WASDx 10 hours ago [-]
Likely because it uses fewer thinking tokens (that you don't see anyways).
algoth1 1 days ago [-]
Just got it. European plus user here. Only codex, no chatgpt
friendlypenguin 19 hours ago [-]
I was really hoping for Astra to be less expensive then Opus...
redox99 18 hours ago [-]
It is (uses way less tokens)
simianwords 18 hours ago [-]
No it isn't cheaper, any source for task vs price comparison to Sol?
This is the second comment I see where you write "no it isn't", without providing a source for your statement. Then you follow it by asking for a source. So is there a source you can provide to back up your statement?
dgellow 17 hours ago [-]
It’s supposed to work the other way, if you make a positive claim that it is cheaper, where is your proof?
ebiester 12 hours ago [-]
For those on plus, are you seeing astra limited to medium? Considering the rate limits that probably makes sense, but I'm wondering what different groups have access to.
d2p 16 hours ago [-]
Odd that the tool call failure rate is so high (5%) for the OpenAI provider than Azure (0.2-0.5%).
NSUserDefaults 12 hours ago [-]
I misread the title as GPTA-6, wondering what sort of crazy crossover was happening.
swe_dima 12 hours ago [-]
According to the metrics the "fast" mode is not any faster...
christophilus 10 hours ago [-]
Tangentially related, but Astra generated some of the worst Odin code I’ve ever seen. Turns out AGI is indistinguishable from an Oracle subcontractor who hates tech and hates his job.
jaesonaras 1 days ago [-]
Anyone had success using Astra as a Foundry model via Github Copilot? The error I get is that tooling is not available if reasoning has a value.
upcoming-sesame 17 hours ago [-]
Any tips on using Astra as orchestrator with Luna workers efficiently in codex?
logged4upvoting 16 hours ago [-]
(This applies to Sol but probably works for Astra too)
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).
DrProtic 13 hours ago [-]
Could you shoot us the github gist?
r_lee 1 days ago [-]
is Azure for this actually ZDR?
jiggawatts 18 hours ago [-]
It cracks me up that you have to ask in a public forum because the vendors purposefully obscure this critical information.
The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.
There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.
If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!
r_lee 15 hours ago [-]
from my understanding, ZDR for this is for vetted customers only + their whole "privacy preserving abuse monitoring" or whatever is active
but again, seems like there's no word from Azure if this applies to them.
very confusing.
1 days ago [-]
azinman2 9 hours ago [-]
I’m confused. I thought it was being only rolled out to select partners for now?
gavinray 1 days ago [-]
I have GPT-6 access in Codex and OpenAI API now
I'm a Business plan user with Cyber verification enabled, FWIW.
embedding-shape 1 days ago [-]
Same just got access literally this minute, Pro user here, no Cyber verification but have passed my ID over to them back in 2024 or something, maybe at the ChatGPT 3 API launch or something?
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
ellessarr 7 hours ago [-]
2× price only wins for review if it catches bugs the cheap model drops — nobody runs that test, everyone quotes the benchmark.
forrestthewoods 17 hours ago [-]
Threw $10 at this to help me prepare for my league’s fantasy auction this weekend. It spend $3.50 and then said “this action would cause you to go above your spending limit”.
Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.
Sure seems like Astra is expensive AF.
starik36 1 days ago [-]
What is the actual utility of using this model on Azure? It's twice as expensive, according to the link.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?
hhh 1 days ago [-]
It's the same price as regular processing. You get guarantees microsoft give you, which are ones OpenAI won't (or require dedicated spend,) and you can use azure identities for access.
claiir 1 days ago [-]
They’re ZDR and the OAI ones aren’t
nibbleyou 20 hours ago [-]
Does anyone know if it is not ZDR on the codex app usage too. I am assuming it isn't
1 days ago [-]
olalonde 24 hours ago [-]
WTHIT?
spdustin 23 hours ago [-]
ZDR = Zero Data Retention — they don't store your inputs/outputs.
viccis 19 hours ago [-]
AKA extortion
DaSHacka 8 hours ago [-]
More like the business tax
Peanuts99 17 hours ago [-]
If you've got access to an Azure subscription and you don't pay the bill personally.
itsjustkev 1 days ago [-]
Compared to OpenAI flex? I'm pretty sure that is their batch processing endpoint, which is naturally cheaper.
Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
https://www.svgviewer.dev/s/i8t1VXzQ
I also think there's an extra level that I would hope an AI would nail which maybe an amateur artist would also fail at, such as thinking about what position a pelican would actually ride a bike in (maybe angling the beak down for aerodynamics etc.), but we are far away from this.
Quote from the thinking trace:
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
Did even better from the front. What's surprising is that it used the exact same colors as the simonw example, despite my prompt only being
> Generate an SVG of a ring-tailed lemur riding an electric scooter. Front view
GPT-6 Astra Extra High
If that was the case, the models would have been producing near perfect outputs for it a year ago.
Instead they are just training on general SVG generation, which in no way should be viewed as "benchmaxxing".
Yes
I'd be shocked if they didn't myself.
Treat it like a bit as is
It feels silly to say that about making a pelican image but it really shows the difference in output and cost in an easy to understand way.
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
Loving the pelican silliness. My ChatGPT is over the moon about it. Gonna miss it when it really is done. Though maybe a pelican riding a bike game could become the benchmark in a year.
I used Light mode and it used up all my limits for the day and had to continue the following day.
Missing GLM-5.3, though.
Reminded me of https://clocks.brianmoore.com/
Coz who knows if astra low will produce max like output if tried once more.
I heard a rumor that Luna is a slightly different architecture from Sol and Terra, which makes me wonder if Luna and Astra might be more related to each other than to Sol and Terra.
Bit of a big leap to make from a token count though!
I just checked the tiktoken library and couldn't see any changes relating to Luna: https://github.com/openai/tiktoken
I'm wondering if this is being trained on by the models today.
https://openai.com/index/advancing-the-price-performance-fro...
Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog
Apparently the crowd agrees because they keep upvoting these.
The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.
Isnt Netherlands the leader in bike riders and they dont wear helmets.
I would recommend staying away from OpenRouter. No matter how good the service is, if anything does go wrong, you have no recourse and you lose every credit in your account. Ironically some of the few responses I actually saw in the Discord were doubling down on their “no refunds no matter what” policy.
Here's an image design source of truth: https://image.non.io/78f4cd8b-2560-4643-9a51-96a89171f994.we...
And here's the page it build from it: https://image.non.io/e7d3a9e5-f9df-4fd8-b79f-1f90280f978f.we...
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.
Everything costs more now than it did a year ago... except for THIS, and we are still complaining that a 90-99% reduction in cost is STILL too expensive. And a 50-75% reduction in time is STILL too long.
We used to have to wait for weeks for a design like that when I worked at a consultancy, and that is a week of salary. For the design, then it got handed off to a front end developer to slice it and get built so the back end developer can hook it up to a CRM. We are talking a month turn around with design, revisions, development, testing, and bug fixing.
It can now be done in a couple of hours for less than a single hour's cost. If it were 10x slower and 10x more expensive, it would STILL be "good deal".
Compared to what?
Yeah if you are solo developer without budget.
For any business this is nothing, the ROI is massive.
It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
It's more "correct" but looks a lot worse in my opinion:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
https://aibenchy.com/showcase/?page=2#showcase=67fc6d6c8e4c3...
https://aibenchy.com/showcase/?page=3#showcase=c215b5c915da6...
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
You can view here all generations for all models: https://aibenchy.com/showcase/
What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all
I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.
I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).
>Playing ping pong from the side of the table
I think marketing might be getting a bit absurd at this point
I even keep seeing obvious stealth marketing like this: "<topic> and how do I use it with <product> in <product>"
The entire mainstream media and political establishment, and every normie I meet is convinced Skynet is already upon us
I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
It must be super interesting working there rn :D
almost had the feelin it was watching its little brother fail and had to 'step in' for a moment :').
time to go play outside...
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
Edit: nevermind it JUST gave me a notification to use it!
It's terse, like all GPT models by default, but the sentences feel less obscurantist.
It's also more pro-active about problem solving.
Where do we go from here??
Edit:
GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens
GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens
So for the same measured intelligence, Sol costs only 40% as much — i.e. ~60% cheaper, while Astra is ~2.5× more expensive.
Why is the burden of proof on me tho!?
astra high is also 3x cheaper than opus max at basically the same intelligence.
astra high is also about as expensive as sol max while being more intelligent.
astra medium is cheaper than sol max while also being cheaper and roughly same intelligence.
im going to replace my sol usage with astra high/medium i think
caveat: benchmarks are really fuzzy with llms
https://artificialanalysis.ai/models/releases/gpt-6-astra
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).
The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.
There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.
If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!
but again, seems like there's no word from Azure if this applies to them.
very confusing.
I'm a Business plan user with Cyber verification enabled, FWIW.
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.
Sure seems like Astra is expensive AF.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?