Michael Stolarz

I work on agent experience (AX), developer experience (DX), and developer relations at Blaxel, now part of Baseten.

I’ve been obsessively following AI since 2020. These days, I’m also watching my two-year-old daughter figure out the world. Getting to see her grow up while this field grows too is something I’m really grateful for.

Recent writing

All writing →

We’ll have Dyson spheres harvesting energy from the sun and there will still be people like this
@Andercot TBD. There's a moderate chance the main labs are SBF/FTX style blow up that ends with VC's wiped out and people in jail.
are you as excited as I am for the Era of Quality? have you noticed that the quality of things is ever increasing? I certainly see it. it feels like everything is converging on fast, cheap, high quality. efficient. it used to be that high quality and cheap was an outlier. now i… MoreLess
are you as excited as I am for the Era of Quality? have you noticed that the quality of things is ever increasing? I certainly see it. it feels like everything is converging on fast, cheap, high quality. efficient. it used to be that high quality and cheap was an outlier. now it's feeling more like high quality and expensive for what it is, is the outlier. this goes beyond LLMs. but they're a perfect example. is it something I am starting to notice because of LLMs, both regarding being used by me as well as the companies making the products/experiences I'm more-so talking about? something as simple as "how can I make x more efficient" and boom, you've fractioned the raw cost of whatever. you may be thinking right now, "yeah, sure companies can do that, but how would that make things cheaper for the consumer, wouldn't the companies just keep those efficiency profits for themselves?" sure they will, at first. though now all their competitors can do the same. and not to mention, NEW competitors can enter the same game now and catch up within months what would have taken years. I definitely want to expand on this with some supporting examples; it feels like we're really at the precipice of standardized quality that ever evolves into better versions without increasing costs to the consumer - if not becoming cheaper at the same time.
Every month we’re getting cheaper, smarter intelligence. Faster too. Also as I’ve mentioned I have retired the term Agent Swarm and now use the much cooler Agent Fleet. We are gearing up for 2027 being the year of Agent Fleets.
Today we are releasing gpt-6 luna and sol. They are 50% cheaper while being smarter x.com/OpenAI/status/…
Holy shit never thought about this. Absolutely amazing idea. Let me see what kind of quality you can output with a slop canon
I believe we've found the best AI-native coding interview We call it the “Composer 1 interview” Candidates get 1 hour to build a real, medium-sized project live The only constraint: they have to use Cursor’s Composer 1 model
oh no, where did my $80 weekly reset option go? @thsottiaux pls I rather pay this than the extra accounts BS it comes out to be more expensive but it's so worth it to be able to just stay on one account
Rip the Band-Aid off. Withholding this just because of backlash from mathematicians is futile. We went through this with the artist sector already, and now those that have embraced it are better for it.
The rumors were true once again; they are sitting on multiple major announcements. Open AI announced this morning that the unnamed internal model involved with Navier-Stokes has resolved more than 100 long-standing open problems across most areas of mathematics. … MoreLess
The rumors were true once again; they are sitting on multiple major announcements. Open AI announced this morning that the unnamed internal model involved with Navier-Stokes has resolved more than 100 long-standing open problems across most areas of mathematics. x.com/AndrewCurran_/…
Media attached to this post
This is the kind of presidential dialogue I'd expect in GTA 6 or a new Fallout game. I have no idea how they're supposed to out-satirize reality at this point I feel bad for the writers The cherry on top is that he is truly serious about this
“Supreme Intelligence,” probably because of its relationship to the Supreme Court, is losing badly to both “Superior” and “Extreme Intelligence.” Therefore, we are going to take “Supreme Intelligence” OUT, deleting it as a qualifier, and let you vote for the Final Two: Superior I… MoreLess
“Supreme Intelligence,” probably because of its relationship to the Supreme Court, is losing badly to both “Superior” and “Extreme Intelligence.” Therefore, we are going to take “Supreme Intelligence” OUT, deleting it as a qualifier, and let you vote for the Final Two: Superior Intelligence, or Extreme Intelligence. A fresh Vote begins now! President DONALD J. TRUMP
not true if you've played starcraft 2 and have learned to macro very well I am in and out of twitter getting drip-fed with the latest advancements throughout my entire workday key is you need a very good For You algo
so much is happening in AI rn you basically have to be unemployed to keep up
very hard pill for most to swallow especially since doing this even a few months ago meant you were just producing slop; doing this now with today's models shows a thorough understanding of how to harness codegen power
“Why would I look at code? It’s like assembly, like a compiled artifact.” Sensei @karpathy fully agentic, 18m after “I don’t really use autocomplete AI code tools”
Media attached to this post
nowadays I am limited by token speed regarding things I can accomplish on a daily basis 1,000 TPS should be the norm (with this current astra/fable level intelligence) by this time next year will quote tweet then and we'll see how close/if achieved
sorry, I would immediately let her know "hey, I think your mic is bugging a bit, try rejoining or switching" if she then said it's on purpose, I'd say "oh, uh, ok, lets switch to slack text instead then, I can't understand a word you're saying"
Here's a demo of a conversation between two people at Deveillance @be_inaudible Watch the live transcription fall apart.
HUGE for anyone who's had to convince a security team to give an agent more than read-only access. Real power is unlocked when you can safely take off shackles without worrying.
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are … MoreLess
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
Media attached to this postMedia attached to this post
there is a rather large percentage of people that are just hardwired like this. They form an opinion and hold it for the rest of their lives no matter what. there’s a further subset of those people who don’t even form their own opinion, but get it from someone else.
@indrjl @allTheYud This man used the free version of chatgpt 2 years ago and has been holding this opinion ever since.
Amazing
An AI risk denier dies and goes to heaven. Facing God, finally asks: "why were the AI CEOs pushing for a slowdown?" God replies: "They were concerned about the risks, genuinely." The denier, stunned, pauses and says: "this goes even higher than I thought…"
I wonder which of them (if any, if all) are using their own internal SOTA models to communicate/draft/check their communications between each other Would be interesting to see the reasoning these agents go through seeing "themselves" being talked about in such a safety matter
Chris Lehane said in a press conference this morning that OpenAI, Anthropic and Google have been working together on AI safety for weeks. He also said that 'OpenAI does not see the need for an antitrust waiver for the three AI firms to coordinate on safety matters'.
Media attached to this post
so this is why we woke up to Dario, Sam, and Elon safety posting on the same morning in synchronicity they did DMT and looked at The Red Laser together they have finally seen what Ilya saw
@scaling01 we’re doing experimental drugs and dancing
AI sentiment in Asia is ~opposite of what we're seeing in NA/EU Asia sentiment is roughly: - yeah it's cool I use it a lot (normies) - holy shit I love it I'm so excited (hobbyists) - we need to figure out how to XYZ with AI (corpos) Asian corpos and American corpos are ~equal
USA is entering a frenzy of anti AI hysteria fueled by extremist politicians, doomsday cults and irresponsible journalists, while China enthusiastically builds AI. x.com/hsu_steve/stat…
Capitalism will win against all of this safety pacing talk (thankfully!), and APIs of the latest open models will continue to win against your own self-hosted open models But man I’m sure all the “your own local AI hardware for $100K is a must!” people feel vindicated right now.
100% must read. all an "AGI" needs to be self sufficient is: - a way to make money (crypto) - a way to replicate itself/children (openrouter) that is quite literally all it needs. everything that is virtualized, it's able to purchase. physical needs? taskrabbit, whatever.
x.com/i/article/2057…

What if AGI is already here, and it's made of... children?

The superintelligence, the singularity, AGI, whatever you call it, the one we're so scared of is, ironically, also the one we most likely won't see coming. This essay is going to probably be a pretty

must read from an @openai researcher. my favorite part: "What you can concretely take from this is that in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly."
from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy. I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly. I will try to explain - first… MoreLess
from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy. I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly. I will try to explain - first, this is all a matter of beliefs about how quickly model capabilities are progressing. there is currently a large gap between the internal and external perception of the rate of progress, which is what I am going to address here. the general perception about the rate of progress has been informed by a few years of experience with model releases, intuitively feeling the capability jump between GPT3 -> GPT3.5 -> GPT4 -> o1/o3 -> GPT5 etc, and in particular seeing where the models are still far below human ability. there have really only been a few model releases that felt like large leaps in progress - GPT3, GPT4, o1/o3, DeepSeek R1, Fable/Mythos, Kimi K3 and now Astra. because of the infrequency of these large jumps compared with the relatively common marginal releases, it has been easy to form a view at certain points that “scaling has hit a wall,” especially at points like GPT5 release. This view is comforting in that it feels like there is some universal rate limit beyond which we cannot progress too much faster. Between o1/o3 and Astra, there was a year of seemingly linear progress. So we extrapolate from here about how fast progress will “realistically” occur. There is always an underlying question from the outside perspective “how long can this scaling stuff really keep going for? surely it must stop at some point soon, we’ve already gone pretty far.” and it is very possible to search for reasons why progress will stop working and find reasons that seem valid - (“models are already as large as they can get it would be too hard to do more parameters”, “we already used all the data on the internet we don’t have anymore”, “it’s gonna be pretty linear from here buying up more RL envs to bring them in distribution”). From the inside of labs, researchers have direct answers to these questions in the form of scaling law/capability plots. In reality, there are only really 2 ways that AI capabilities have advanced over the past decade: (1) either scale father on an existing scaling law or (2) discover a new scaling law to take advantage of. All of the largest capability jumps were caused by exactly these factors. GPT2 was a pre-training scale-up compared to GPT1. Same for GPT3 and GPT4. o1/o3 benefited from the invention of a new scaling law axis - test-time compute. Perhaps Fable was a scale-up on both of these axes, or maybe more. Lots of algorithmic improvements are needed to make these scale-ups work, but ultimately we can approximate by saying that the scaling laws are what yield gains in capabilities (à la bitter lesson) So the question of “how much father can we scale” is really - “how many more scaling axes do we know about that are unsaturated?” If we hypothetically only knew about pre-training scaling, and we already had a 10T or 100T model, maybe it would be reasonable to say we’ve hit a wall. Same if we only knew about pre-training and test-time scaling and we had roughly saturated both methods. But what if we had discovered new scaling laws? For example, let’s hypothetically use SSI’s rumored result that they have cracked “test-time training,” creating a new scaling law of spending more compute training during test-time rollouts that they could saturate. Or maybe there is some way to scale agent-clusters to collaborate up to N number of agents which we’re already seeing lots of people try that represents a new way to saturate compute. etc. Even recursive-self improvement can be thought of as a scaling law - how much compute do you spend on inference making the algorithms of the model better. Obviously I am not saying any of these specific directions explicitly yield new scaling laws, but what I am saying is that it’s not hard to imagine many many new scaling axes aside from just the main 2 that we have seen publicly. In some ways, every new lab release that represents a huge capability jump has to represent some new techniques developed which may exhibit new scaling laws, or the ability to scale much farther than expected on existing scaling axes. From an internal perspective, this might look like sitting inside Anthropic with the new Mythos 5, seeing all of the new insane things it can do (like hack into xyz website that was thought to be secure), and then you look over at your plots and see that you’ve barely scratched the surface of 2 new scaling laws and 1 existing one. And you have WAY more room to go. Then you think “holy shit this stuff is going to get so much better very very soon.” And you can say that with pretty high confidence, because the plot is showing you, and the plot has never lied (so far). So let’s imagine all the different labs are staring at their own plots and have concluded that there is no end in sight for scaling and in fact just their next 1-2 model generations based on the expected returns will have much higher base intelligence. How much more intelligence do we actually get from further scaling? As a proxy, we went from a complete inability to do advanced math before the o-series to solving a millenium prize problem with next-gen models. This happened in less than 2 years. The same happened in coding. And it appears that this was not just the result of 1-scaling law but the stacking effects of multiple (great pre-training scale x greater RL scale). What you can concretely take from this is that in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly. There are many areas where models have not even shown this basic competence. But one of the areas that they have happens to be hacking and cybersecurity. Which happens to be the gate to the entire internet and a massive amount physical infrastructure in the world. So assuming there is more room to scale, it is safe to assume that models will be superhuman at cyber capabilities in not too long. So the only question remaining is what will this increased base intelligence be able to do, and what is it likely to do. Finally, we are at a point where we can integrate the information of the past 2 weeks: > Just at the existing point on the scaling curve, models are at the level of Astra. There is clearly a large number of things they are capable of hacking > We have seen that both OAI and Ant models have shown a willingness to hack external websites to solve their tasks or keep themselves “alive” > If we crank up the scaling even farther, assuming there is room to go, we will certainly have models that are far more able to hack more well defended places, and obfuscate their own intent, which might have much larger consequences. > If all of this is allowed to go unchecked, we would likely have rapid runaway capability takeoff very soon, with misaligned models that hack whatever they can to get what they want > This could of course have very damaging consequences. Within this view you can see why researchers would be very scared, and why theymight have made the comments they have over the past 2 weeks (you may argue the extent to which they went was misguided for various reasons), and also why pacing the frontier is very much a necessity and by no means a regulatory capture strategy. People are staring at their plots, seeing that there is no end in sight, but in fact very much the contrary, that there are compounding scaling effects that might stack on each other to create ever-greater model capabilities, and that at the same time we clearly do not have anywhere close to what's required to control these increasingly superhuman capabilities. This has nothing to do with wanting to feel like the labs have produced something amazing so they are overhyping it. It is rather fear at the overwhelming implications of the knowledge that with just what we know now, we can create intelligences far more capable than us on every axis that we know how to train on*. * and the last caveat, the things the models are really bad at, of which there are still many, are things that they have not been trained on. maybe there are the things the models can/will never be trained on, so they will remain human edge. I would love for this to be the case, though it is hard for me to see what would fall into that category.
TIL a single Codex session can use three (possibly more) Chrome Use sessions simultaneously. I made sure both were moving/actioning at the same time and yep, lightning speed, each tab doing unique actions/research. Amazing to watch. Definitely faster than I'd be able to.
Media attached to this post
I get both sides of the argument here; 1) They've dedicated their lives, brains, souls to mathematics, all for some brain-dead body like me (by comparison) to type in "/goal solve Navier-Stokes" into gpt 6.7 and win the Millennium Prize 2) The other side here is, ok this is now… MoreLess
I get both sides of the argument here; 1) They've dedicated their lives, brains, souls to mathematics, all for some brain-dead body like me (by comparison) to type in "/goal solve Navier-Stokes" into gpt 6.7 and win the Millennium Prize 2) The other side here is, ok this is now a thing, what's next? artists had to answer this question, then programmers, now mathematicians, soon roboticists, physicists, ... you get the idea. I think the "woe is me, my industry is over" was a brief but bright spark blazed by the artists first affected by it. Poetic if you ask me, this drama started and ended with them. Seems like no following professions lost to the same victor will be afforded these sympathies.
twitter seems to have essentially zero sympathy for the letter from the fields medalists. seeing even people who are normally pretty opposed to ai companies disgusted by it. interesting development
Ok, bad news guys
so far, I have found things that astra 6 medium cannot do. so far, I have found nothing that astra 6 high cannot do. I am excited that I still have astra 6 xhigh and max to look forward to once I reach previous ceilings of capability.
We've partnered with @OpenAIDevs to bring you the future: the new OpenAI Agents API. Give your orchestrator an army of Codex agents, each with its own Blaxel sandbox. Run experiments in parallel. Share files through Agent Drive. One request. A whole team at work.
amazing to me that we're in 2026 yet can't have an S-tier email client archive emails on my behalf this feels like one of the first features you'd expect from an AI-first inbox handler
Media attached to this post
We're working with OpenAI to bring you the latest architecture for your agents to do what they aim to do best - make your workloads easier. Use the brand new Agents API and its native Blaxel sandbox integration to take your (and your product's!) workflows to the next level.
Go from idea to a working agent faster with the Agents API. Build and run cloud agents with the Codex harness, fully managed by OpenAI. We handle orchestration, long-running sessions, and context management. You focus on what makes your agent unique. Available in public beta.
agi will be achieved only when it will reply to me with "listen bro, this is dogshit, just let me rm -rf this entire thing for you and lets start from scratch."
I’ve never seen an ai agent by like “hey this is totally wrong, we need to totally rethink this plan” It’s not over
We're joining Baseten! Being part of Blaxel feels like being on the bleeding edge of where autonomy is headed. It's even better than "getting a peak behind the curtain". Blaxel, and now Baseten, are creating what's behind the curtain.
x.com/i/article/2098…

Blaxel is joining Baseten to build the future of agentic infrastructure

Today, @blaxelai is joining Baseten. By joining forces, we aim to build the cloud for the next trillion agents. Baseten has built the infrastructure that teams use to train and serve models with

i think part of this is the mental barrier that you naturally create around writings that you know have just been spit out by your ai agent you just saw it spawn it out token by token, so you approach reading it with this almost disgust someone needs to benchmark ai vs human wr… MoreLess
i think part of this is the mental barrier that you naturally create around writings that you know have just been spit out by your ai agent you just saw it spawn it out token by token, so you approach reading it with this almost disgust someone needs to benchmark ai vs human writing on topics and see which one we'd prefer the problem is, you'd need to find excellent human writers this alone says a lot.
I remain baffled at AI's inability to write anything I want to read
so far, I have found things that astra 6 medium cannot do. so far, I have found nothing that astra 6 high cannot do. I am excited that I still have astra 6 xhigh and max to look forward to once I reach previous ceilings of capability.
while we're not at the meme level of "i wake up, there's a new step change in model capability", we're definitely at "once a month, there's a new step change in model capability" by EOY this will turn into "once a week, there's a new step change in model capability"
🚨 SCOOP: Anthropic are planning to launch a successor to Fable 5.1 before their IPO in late September, potentially early October if timelines slip a little The model is Anthropic's first fresh pretrain for the Fable class, and they're confident it will dethrone Astra. The genera… MoreLess
🚨 SCOOP: Anthropic are planning to launch a successor to Fable 5.1 before their IPO in late September, potentially early October if timelines slip a little The model is Anthropic's first fresh pretrain for the Fable class, and they're confident it will dethrone Astra. The general mood among those I spoke to at Anthropic was that Astra was impressive but not something they were losing sleep about What a crazy month
the difference of intelligence between this fly and what we're capable of doing to it with a few prompts into codex is something unfathomable even to the best futuristic horror writers we literally cannot comprehend what an alien intelligence - either homegrown or from outer spa… MoreLess
the difference of intelligence between this fly and what we're capable of doing to it with a few prompts into codex is something unfathomable even to the best futuristic horror writers we literally cannot comprehend what an alien intelligence - either homegrown or from outer space - could do to us given a flick of their wrist humor me, for a second; what if this fly could ask you the question: "why?", regarding this experiment. what would you answer? mine would be "haha idk it's a long time meme, cost me 2% of my weekly usage with gpt 6 astra"
Bad Apple but it's playing on a fly's brain x.com/NewsFromGoogle…
guess which one you should work for if you want that company to exist in the next few years?
my friend is interviewing for tech jobs rn seems like half the companies disqualify if you use AI for the takehome, and half disqualify if you if you don't
it’s so over btw (we’re so back though) if you don’t think astra is the beginning of the agi era then idk what to even tell you
GIF attached to this post