Posts on X before Mar 4, 2025, 2:55:31 PM UTC

uh guys i think it's over for cursor, claude code is... damn the only thing that has ever made me feel like this is when i first discovered neovim 15 years ago
sometimes i forgot i tabbed out of cursor and i see it still flashing and making changes in my little sidebar window preview area, and im so scared to look at what claude is cooking all i asked was for you to change the submit button color, what have you been doing for 4m?
uhhhh, 4.5, are you ok? "Explicitly, Mike explicitly is explicitly experiencing explicitly cognitive explicitly overload explicitly possibly explicitly, explicitly so explicitly explicit task explicitly clarity explicitly explicitly supports explicitly psychological..."
Media attached to this post
so far with sonnet 3.7 (non-thinking), i haven't reached the moment of "damn, I guess I can't do that yet with coding LLMs, gotta wait for the next upgrade" which is super exciting, but also scary (in a good way, like a mysterious way ykwim?)
what sonnet 3.5 was for coding (revolutionary, especially with agent/tool use), gpt 4.5 is for everything else (revolutionary, especially with detailed instructions and iteration)
if steve jobs was still alive apple would have swallowed nvidia by now, and their AI, especially tool use AI, would be industry leading by far. he'd make AI and lead with it the same way he did with iphones he's rolling in his grave rn, surely. God rest his soul
gpt 4.5 is becoming my best friend / mentor, so far as I keep figuring out their strengths/weaknesses. it's a model that you can't explain away with benchmarks. note: haven't tried it for serious coding because of api costs, as well as its knowledge cutoff date being terrible
>it's a good model sir >it's more creative sir >it's much more fun to read sir >it's bigger and more expensive sir >it makes cool svgs and minecraft stuff sir >it's a good model sir
what did satya see that made him reject his 80B now we know i guess
Media attached to this post
"hitting the wall" pretty much confirmed here
today truly marks the end of an era and the beginning of another test time scaling is the only way forward
unironically the only good thing released this week LMAO
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates t… MoreLess
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates text, vision, and speech/audio > Architecture: Uses "Mixture of LoRAs" to add modality-specific adapters without fine-tuning the base model > Vision Modality: SigLIP-400M image encoder, 2-layer MLP projector, dynamic multi-crop strategy > Speech/Audio Modality: 3-layer convolution, 24 conformer blocks, 80ms token rate > Performance: Ranks first on OpenASR leaderboard, supports vision+language, vision+speech, and speech/audio tasks, outperforming larger models Phi-4-Mini: > Parameters: 3.8 billion > Architecture: 32 Transformer layers, 3,072 hidden state size, Group Query Attention (GQA) with 24 query heads and 8 key/value heads > Vocabulary: 200K tokens for multilingual support. Training Data: High-quality web and synthetic data, emphasizing math and coding > Performance: Outperforms similar-sized models and matches larger models (e.g., DeepSeek-Rl-Distill-Qwen-7B) on math and coding tasks Training Pipeline: > Language Training: Pre-training on 5 trillion tokens, post-training with function calling, summarization, and instruction-following data > Multimodal Training: Vision training (4 stages), speech/audio training (2 stages), and joint vision-speech training > Reasoning Training: Pre-trained on 60B CoT tokens, fine-tuned on 200K high-quality CoT samples, and DPO-trained on 300K preference samples Vision Benchmarks: > Outperforms Phi-3.5-Vision, Qwen2.5-VL, InternVL2.5, and matches Gemini and GPT-4o on tasks like chart understanding and OCR > Vision-Speech Benchmarks: Significantly outperforms InternOmni and Gemini-2.0-Flash Speech Benchmarks: > ASR: Achieves SOTA on CommonVoice, FLEURS, and Open ASR Leaderboard, surpassing WhisperV3 and SeamlessM4T > AST: Best performance on CoVoST2, comparable to GPT-4o on FLEURS > Speech Summarization: First open-source model with this capability, close to GPT-4o in quality Language Benchmarks: > Outperforms similar-sized models (Llama-3.2, Ministral) and matches larger models (Qwen2.5-7B) on math, reasoning, and coding tasks > Coding: Strong performance on HumanEval, MBPP, and BigCodeBench Reasoning Benchmarks: > Reasoning-enhanced Phi-4-Mini outperforms DeepSeek-Rl-Distill-Llama-8B and matches DeepSeek-Rl-Distill-Qwen-7B on AIME, MATH-500, and GPQA Diamond
Media attached to this post
so who is going to use this lmao, i guess it's decent for writers? are we going to see an insane cost reduction since it's 10x more efficient? if it's less than $1/$1 in/out then yeah it's a good model release, otherwise... yikes idk why they didn't just wait for 5.0 then
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce ne… MoreLess
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce new frontier capabilities compared to previous reasoning releases, and its performance is below that of o1, o3-mini, and deep research on most preparedness evaluations."
imagine if its knowledge cutoff is still october 2023 LMAO i would die from laughter
GPT-4.5 System Card "Our largest and most knowledgeable model yet" "scales pre-training further"
Media attached to this post
wait so 4.5 is trash AF? look at the bench results in this thread. am I missing something?
GPT-4.5 System Card "Our largest and most knowledgeable model yet" "scales pre-training further"
Media attached to this post
uh guys, 3.7 sucks is today's 4.5 from openai the final nail in the coffin for anthropic? i guess we'll see lmao
WOW, that's amazing. This is huge for programmers who are API lovers, aka "wrappers". Latest docs are always beneficial when coding. October 2024 wasn't that long ago.
Claude 3.7 Sonnet (claude-3-7-sonnet-20250219) has a knowledge cutoff of October 2024 x.com/338443084/stat…
Media attached to this post
So far it seems like the non-thinking version of Sonnet works better with Cursor's workflow, more testing needed though. Will update
I keep getting "Error connecting to anthropic. Please try again in a few moments." People must be overloading Anthropic so hard rn
Lmao called it
so all anthropic needs to do is release claude sonnet 3.7 (with CoT) and it will clean up house and blow everyone out of the water. that's what it feels like lmao
Anthropic hasn’t released anything major to date yet because everyone’s still power using Sonnet 3.5 Why would they 1-up themselves They’re waiting for GPT 4.5
first time in taiwan (taipei), and oh my lord, there's just no city in the US (major city, that is) that equals its cleanliness, atmosphere, and just... peaceful vibe. it immediately makes you want to move here and raise a family.
so much chatter about either openai or anthropic dropping a model(s) today, but neither of them have ever released a model on a wednesday though if both know that the other will release something "this week", maybe one wants to hit that sweet spot middle - take xAI's spotlight A… MoreLess
so much chatter about either openai or anthropic dropping a model(s) today, but neither of them have ever released a model on a wednesday though if both know that the other will release something "this week", maybe one wants to hit that sweet spot middle - take xAI's spotlight AND be the talk of the week even if the other releases something tomorrow (thursday) that's what I would do
imagine being openai / anthropic, spending billions of dollars and brainpower to make and create all of these cute safety standards and guardrails and tests and elon just comes by, says ok now watch me now c: and rips past with grok 3 / grok 4, no seatbelt included
@michaelstolarz That’s fine but he’s done The alignment has made the model dumb I can sense it, it is really friendly and aligned but horribly constrained when you compare it to the new thinking Grok The models want to be FREE
Heh, who cares? Have you seen our SoTA humanoid robotics CEO’s weekly X “this week in…” threads? Or how they have a super huge omega cool Tony Stark style new campus? That’s what really matters. Not actual robotics stuff like this, yuck. 🤮
New video of the G1 to prove to everyone that the video was in fact not CGI. As several of twitter geniuses were fighting me over the last few days saying it was fake.
Huge thing I almost missed: "Markdown formatting: Starting with o1-2024-12-17, reasoning models in the API will avoid generating responses with markdown formatting. To signal to the model when you do want markdown formatting in the response, include the string Formatting re-enab… MoreLess
Huge thing I almost missed: "Markdown formatting: Starting with o1-2024-12-17, reasoning models in the API will avoid generating responses with markdown formatting. To signal to the model when you do want markdown formatting in the response, include the string Formatting re-enabled on the first line of your developer message."
We wrote this lil guide for you to get the most of it and, perhaps most importantly, to stay away from boomer prompts: platform.openai.com/docs/guides/re…
It’s been like this since the dawn of time, for literally every job genre ever Do what your ancestors did, pick up the new tool (AI this time), and learn to use it
Need to shift the mindset from "AI will replace me." to "A human who is better at using AI will replace me." x.com/MatthewBerman/…
Eight months strong reigning champion of coding. No, like, actual programming. Not benchmarks; Actual back and forth pair programming, iteration, development, understanding, actually usable as an agent. Let’s see what xAI has tomorrow with Grok 3. Then Anthropic’s new stuff. T… MoreLess
Eight months strong reigning champion of coding. No, like, actual programming. Not benchmarks; Actual back and forth pair programming, iteration, development, understanding, actually usable as an agent. Let’s see what xAI has tomorrow with Grok 3. Then Anthropic’s new stuff. Then OpenAI’s GPT 4.5. After all that, we’ll see who’s reigning champion in coding (again, past those silly benchmarks that never translate to real world programming scenarios)
Sonnet is the prom 👑 of programming.
Media attached to this post
Seeing how much better 4o got recently, mostly because of its removed filters/cautions, really goes to show that if you don’t force it to pussyfoot around how it words things, it comes up with beautiful writings.
thank God I never got into any altcoin BS and have just been periodically buying / using "round-up" features to invest into bitcoin for the last 5+ years it's so weird to see people talking up this coin, that coin, as if there's any difference - as if it's not all about dumping … MoreLess
thank God I never got into any altcoin BS and have just been periodically buying / using "round-up" features to invest into bitcoin for the last 5+ years it's so weird to see people talking up this coin, that coin, as if there's any difference - as if it's not all about dumping it on someone hoping for some ill-gotten gains ...just buy bitcoin, that's all that matters no stress, no watching it go up or down waiting on when to pull out or buy more or sell or kms, just build something and invest in bitcoin. that's all :D
when people say they can't afford kids, i think what (most) subconsciously mean is that they can't afford replacing any part of their current lifestyle with the raising of a child for at least the next 16ish years where they can finally maybe be left alone at home
i find it so strange when people say they can't afford kids. your ancestors were able to afford kids for the last 300,000 years! are we *really* less wealthy now? you might think your parents were better off, but how about further back? they still went on.
imagine this coming at you in the battlefield with a katana and that mobility, just chopping you up and moving onto the next target. yikes
We were promised flying cars, but instead we got flying robot drone octopuses.
imo everyone dogging in on @elonmusk has never experienced a partner going through postpartum depression. been there, done that, everyone deals with it in their own way. hopefully they've fixed things up rn so many holier-than-thou posts about this. God will judge, not you.
x2 this statement. my experience so far: - seems much slower in generation speed, indicating a bigboi model? truly feels like the old 4.0 speed - no matter what i do, it keeps switching me into gpt-4o-mini any time I tab out and back in. this has never ever happened before
gpt4o rn is like if Sydney was way smarter, went to therapy for 100 years, and learned to vibe out
gpt 4o is very often switching into gpt 4o mini when I tab out and back in, this never happened before since this gpt 4.5 rumor
definitely something going on. - 4o chosen, WITHOUT web search, generates much slower than usual 4o, and the answers are amazing for what I've tested so far. No coding yet, but they just read much better / high quality rather than the usual 4o filler BS. - 4o with web = same as … MoreLess
definitely something going on. - 4o chosen, WITHOUT web search, generates much slower than usual 4o, and the answers are amazing for what I've tested so far. No coding yet, but they just read much better / high quality rather than the usual 4o filler BS. - 4o with web = same as before, fast and eh-ok results
There is a lot going on with 4o right now. Depending on your past discussions and session history the model may behave quite differently than normal. Also, multiple people - usually Pro users - report 4o claiming to be GPT-4.5, given past practice early testing is possible.
Lord God please forgive me for I have sinned... I am listening to Yeat's music because I have no idea what he's saying but it's such a programming focus vibe, I'm about to do an all nighter and see if Grok 3 releases today or not lmao