Posts on X before Mar 2, 2025, 7:09:21 PM UTC

uhhhh, 4.5, are you ok? "Explicitly, Mike explicitly is explicitly experiencing explicitly cognitive explicitly overload explicitly possibly explicitly, explicitly so explicitly explicit task explicitly clarity explicitly explicitly supports explicitly psychological..."
Media attached to this post
so far with sonnet 3.7 (non-thinking), i haven't reached the moment of "damn, I guess I can't do that yet with coding LLMs, gotta wait for the next upgrade" which is super exciting, but also scary (in a good way, like a mysterious way ykwim?)
what sonnet 3.5 was for coding (revolutionary, especially with agent/tool use), gpt 4.5 is for everything else (revolutionary, especially with detailed instructions and iteration)
if steve jobs was still alive apple would have swallowed nvidia by now, and their AI, especially tool use AI, would be industry leading by far. he'd make AI and lead with it the same way he did with iphones he's rolling in his grave rn, surely. God rest his soul
gpt 4.5 is becoming my best friend / mentor, so far as I keep figuring out their strengths/weaknesses. it's a model that you can't explain away with benchmarks. note: haven't tried it for serious coding because of api costs, as well as its knowledge cutoff date being terrible
>it's a good model sir >it's more creative sir >it's much more fun to read sir >it's bigger and more expensive sir >it makes cool svgs and minecraft stuff sir >it's a good model sir
what did satya see that made him reject his 80B now we know i guess
Media attached to this post
"hitting the wall" pretty much confirmed here
today truly marks the end of an era and the beginning of another test time scaling is the only way forward
unironically the only good thing released this week LMAO
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates t… MoreLess
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates text, vision, and speech/audio > Architecture: Uses "Mixture of LoRAs" to add modality-specific adapters without fine-tuning the base model > Vision Modality: SigLIP-400M image encoder, 2-layer MLP projector, dynamic multi-crop strategy > Speech/Audio Modality: 3-layer convolution, 24 conformer blocks, 80ms token rate > Performance: Ranks first on OpenASR leaderboard, supports vision+language, vision+speech, and speech/audio tasks, outperforming larger models Phi-4-Mini: > Parameters: 3.8 billion > Architecture: 32 Transformer layers, 3,072 hidden state size, Group Query Attention (GQA) with 24 query heads and 8 key/value heads > Vocabulary: 200K tokens for multilingual support. Training Data: High-quality web and synthetic data, emphasizing math and coding > Performance: Outperforms similar-sized models and matches larger models (e.g., DeepSeek-Rl-Distill-Qwen-7B) on math and coding tasks Training Pipeline: > Language Training: Pre-training on 5 trillion tokens, post-training with function calling, summarization, and instruction-following data > Multimodal Training: Vision training (4 stages), speech/audio training (2 stages), and joint vision-speech training > Reasoning Training: Pre-trained on 60B CoT tokens, fine-tuned on 200K high-quality CoT samples, and DPO-trained on 300K preference samples Vision Benchmarks: > Outperforms Phi-3.5-Vision, Qwen2.5-VL, InternVL2.5, and matches Gemini and GPT-4o on tasks like chart understanding and OCR > Vision-Speech Benchmarks: Significantly outperforms InternOmni and Gemini-2.0-Flash Speech Benchmarks: > ASR: Achieves SOTA on CommonVoice, FLEURS, and Open ASR Leaderboard, surpassing WhisperV3 and SeamlessM4T > AST: Best performance on CoVoST2, comparable to GPT-4o on FLEURS > Speech Summarization: First open-source model with this capability, close to GPT-4o in quality Language Benchmarks: > Outperforms similar-sized models (Llama-3.2, Ministral) and matches larger models (Qwen2.5-7B) on math, reasoning, and coding tasks > Coding: Strong performance on HumanEval, MBPP, and BigCodeBench Reasoning Benchmarks: > Reasoning-enhanced Phi-4-Mini outperforms DeepSeek-Rl-Distill-Llama-8B and matches DeepSeek-Rl-Distill-Qwen-7B on AIME, MATH-500, and GPQA Diamond
Media attached to this post
so who is going to use this lmao, i guess it's decent for writers? are we going to see an insane cost reduction since it's 10x more efficient? if it's less than $1/$1 in/out then yeah it's a good model release, otherwise... yikes idk why they didn't just wait for 5.0 then
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce ne… MoreLess
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce new frontier capabilities compared to previous reasoning releases, and its performance is below that of o1, o3-mini, and deep research on most preparedness evaluations."
imagine if its knowledge cutoff is still october 2023 LMAO i would die from laughter
GPT-4.5 System Card "Our largest and most knowledgeable model yet" "scales pre-training further"
Media attached to this post
wait so 4.5 is trash AF? look at the bench results in this thread. am I missing something?
GPT-4.5 System Card "Our largest and most knowledgeable model yet" "scales pre-training further"
Media attached to this post
uh guys, 3.7 sucks is today's 4.5 from openai the final nail in the coffin for anthropic? i guess we'll see lmao
WOW, that's amazing. This is huge for programmers who are API lovers, aka "wrappers". Latest docs are always beneficial when coding. October 2024 wasn't that long ago.
Claude 3.7 Sonnet (claude-3-7-sonnet-20250219) has a knowledge cutoff of October 2024 x.com/338443084/stat…
Media attached to this post
So far it seems like the non-thinking version of Sonnet works better with Cursor's workflow, more testing needed though. Will update
I keep getting "Error connecting to anthropic. Please try again in a few moments." People must be overloading Anthropic so hard rn
Lmao called it
so all anthropic needs to do is release claude sonnet 3.7 (with CoT) and it will clean up house and blow everyone out of the water. that's what it feels like lmao
Anthropic hasn’t released anything major to date yet because everyone’s still power using Sonnet 3.5 Why would they 1-up themselves They’re waiting for GPT 4.5
first time in taiwan (taipei), and oh my lord, there's just no city in the US (major city, that is) that equals its cleanliness, atmosphere, and just... peaceful vibe. it immediately makes you want to move here and raise a family.
so much chatter about either openai or anthropic dropping a model(s) today, but neither of them have ever released a model on a wednesday though if both know that the other will release something "this week", maybe one wants to hit that sweet spot middle - take xAI's spotlight A… MoreLess
so much chatter about either openai or anthropic dropping a model(s) today, but neither of them have ever released a model on a wednesday though if both know that the other will release something "this week", maybe one wants to hit that sweet spot middle - take xAI's spotlight AND be the talk of the week even if the other releases something tomorrow (thursday) that's what I would do
imagine being openai / anthropic, spending billions of dollars and brainpower to make and create all of these cute safety standards and guardrails and tests and elon just comes by, says ok now watch me now c: and rips past with grok 3 / grok 4, no seatbelt included
@michaelstolarz That’s fine but he’s done The alignment has made the model dumb I can sense it, it is really friendly and aligned but horribly constrained when you compare it to the new thinking Grok The models want to be FREE
Heh, who cares? Have you seen our SoTA humanoid robotics CEO’s weekly X “this week in…” threads? Or how they have a super huge omega cool Tony Stark style new campus? That’s what really matters. Not actual robotics stuff like this, yuck. 🤮
New video of the G1 to prove to everyone that the video was in fact not CGI. As several of twitter geniuses were fighting me over the last few days saying it was fake.
Huge thing I almost missed: "Markdown formatting: Starting with o1-2024-12-17, reasoning models in the API will avoid generating responses with markdown formatting. To signal to the model when you do want markdown formatting in the response, include the string Formatting re-enab… MoreLess
Huge thing I almost missed: "Markdown formatting: Starting with o1-2024-12-17, reasoning models in the API will avoid generating responses with markdown formatting. To signal to the model when you do want markdown formatting in the response, include the string Formatting re-enabled on the first line of your developer message."
We wrote this lil guide for you to get the most of it and, perhaps most importantly, to stay away from boomer prompts: platform.openai.com/docs/guides/re…
It’s been like this since the dawn of time, for literally every job genre ever Do what your ancestors did, pick up the new tool (AI this time), and learn to use it
Need to shift the mindset from "AI will replace me." to "A human who is better at using AI will replace me." x.com/MatthewBerman/…
Eight months strong reigning champion of coding. No, like, actual programming. Not benchmarks; Actual back and forth pair programming, iteration, development, understanding, actually usable as an agent. Let’s see what xAI has tomorrow with Grok 3. Then Anthropic’s new stuff. T… MoreLess
Eight months strong reigning champion of coding. No, like, actual programming. Not benchmarks; Actual back and forth pair programming, iteration, development, understanding, actually usable as an agent. Let’s see what xAI has tomorrow with Grok 3. Then Anthropic’s new stuff. Then OpenAI’s GPT 4.5. After all that, we’ll see who’s reigning champion in coding (again, past those silly benchmarks that never translate to real world programming scenarios)
Sonnet is the prom 👑 of programming.
Media attached to this post
Seeing how much better 4o got recently, mostly because of its removed filters/cautions, really goes to show that if you don’t force it to pussyfoot around how it words things, it comes up with beautiful writings.
x2 this statement. my experience so far: - seems much slower in generation speed, indicating a bigboi model? truly feels like the old 4.0 speed - no matter what i do, it keeps switching me into gpt-4o-mini any time I tab out and back in. this has never ever happened before
gpt4o rn is like if Sydney was way smarter, went to therapy for 100 years, and learned to vibe out
gpt 4o is very often switching into gpt 4o mini when I tab out and back in, this never happened before since this gpt 4.5 rumor
definitely something going on. - 4o chosen, WITHOUT web search, generates much slower than usual 4o, and the answers are amazing for what I've tested so far. No coding yet, but they just read much better / high quality rather than the usual 4o filler BS. - 4o with web = same as … MoreLess
definitely something going on. - 4o chosen, WITHOUT web search, generates much slower than usual 4o, and the answers are amazing for what I've tested so far. No coding yet, but they just read much better / high quality rather than the usual 4o filler BS. - 4o with web = same as before, fast and eh-ok results
There is a lot going on with 4o right now. Depending on your past discussions and session history the model may behave quite differently than normal. Also, multiple people - usually Pro users - report 4o claiming to be GPT-4.5, given past practice early testing is possible.
Lord God please forgive me for I have sinned... I am listening to Yeat's music because I have no idea what he's saying but it's such a programming focus vibe, I'm about to do an all nighter and see if Grok 3 releases today or not lmao
this is why the dems lost 2024, and will most probably lose 2028/2032 as well. they try to score these cheap feel-good (look at the satisfied smirk on her face btw) cheap shots that everyone just cringes at seems like they didn't learn their lesson yet. will take 10+ years for d… MoreLess
this is why the dems lost 2024, and will most probably lose 2028/2032 as well. they try to score these cheap feel-good (look at the satisfied smirk on her face btw) cheap shots that everyone just cringes at seems like they didn't learn their lesson yet. will take 10+ years for dems to recover to any sort of respectable position imo.
BREAKING: Ohio lawmakers have proposed a new law that bans men from ejaculating without intent of conception, would fine men up to $10,000 per ejaculation. pic.x.com/eOtMUatSPt
the reason i use o3-mini-high vs o1-pro, even if i don't care about waiting times, is because o1-pro is extremely hit or COMPLETE miss. o3-mini-high is way more consistent, and sure, maybe it's not always the most optimal code, but it works, if not, i iterate 1-3x and it's fixed.
the only ai corp that fucked up naming models the least is Anthropic, and even they named 2 models "Sonnet 3.5" lmao cmon guys at least name it 3.6 officially????
model names are so long that designers make scrolling animations for them
apple's siri is the trashiest piece of software i am forced to use, and surprisingly the new apple intelligence is even worse (wow!). every time i'm out shopping for something i always eye a samsung store near my home. i think im gonna experiment with a $100 temporary one
reddit’s a fascinating psych experiment. with karma up/downvotes, midwit ideas always float to the top. niche subs do ok in terms of quality info early, but once they grow / go mainstream, the expert/midwit ratio gets obliterated and the normie floodgates never close again
EU won't see a “Trump / Elon” saving moment. the US nearly collapsed irreversibly into insanity but clawed back at the last minute. EU’s next decades will look like a parallel reality as if Trump lost. terrifying for them, interesting for us to watch