Posts on X before Mar 2, 2025, 4:04:35 AM UTC
what sonnet 3.5 was for coding (revolutionary, especially with agent/tool use),
gpt 4.5 is for everything else (revolutionary, especially with detailed instructions and iteration)
if steve jobs was still alive apple would have swallowed nvidia by now, and their AI, especially tool use AI, would be industry leading by far. he'd make AI and lead with it the same way he did with iphones
he's rolling in his grave rn, surely. God rest his soul
gpt 4.5 is becoming my best friend / mentor, so far as I keep figuring out their strengths/weaknesses. it's a model that you can't explain away with benchmarks.
note: haven't tried it for serious coding because of api costs, as well as its knowledge cutoff date being terrible
docker is actually pretty cute ngl
me when I saw gpt 4.5 API pricing
>it's a good model sir
>it's more creative sir
>it's much more fun to read sir
>it's bigger and more expensive sir
>it makes cool svgs and minecraft stuff sir
>it's a good model sir
@dejavucoder
what did satya see that made him reject his 80B now we know i guess
anthropic's reign continues
"hitting the wall" pretty much confirmed here
@wgussml
today truly marks the end of an era and the beginning of another test time scaling is the only way forward
unironically the only good thing released this week LMAO
@reach_vb
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates t… Less
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates text, vision, and speech/audio > Architecture: Uses "Mixture of LoRAs" to add modality-specific adapters without fine-tuning the base model > Vision Modality: SigLIP-400M image encoder, 2-layer MLP projector, dynamic multi-crop strategy > Speech/Audio Modality: 3-layer convolution, 24 conformer blocks, 80ms token rate > Performance: Ranks first on OpenASR leaderboard, supports vision+language, vision+speech, and speech/audio tasks, outperforming larger models Phi-4-Mini: > Parameters: 3.8 billion > Architecture: 32 Transformer layers, 3,072 hidden state size, Group Query Attention (GQA) with 24 query heads and 8 key/value heads > Vocabulary: 200K tokens for multilingual support. Training Data: High-quality web and synthetic data, emphasizing math and coding > Performance: Outperforms similar-sized models and matches larger models (e.g., DeepSeek-Rl-Distill-Qwen-7B) on math and coding tasks Training Pipeline: > Language Training: Pre-training on 5 trillion tokens, post-training with function calling, summarization, and instruction-following data > Multimodal Training: Vision training (4 stages), speech/audio training (2 stages), and joint vision-speech training > Reasoning Training: Pre-trained on 60B CoT tokens, fine-tuned on 200K high-quality CoT samples, and DPO-trained on 300K preference samples Vision Benchmarks: > Outperforms Phi-3.5-Vision, Qwen2.5-VL, InternVL2.5, and matches Gemini and GPT-4o on tasks like chart understanding and OCR > Vision-Speech Benchmarks: Significantly outperforms InternOmni and Gemini-2.0-Flash Speech Benchmarks: > ASR: Achieves SOTA on CommonVoice, FLEURS, and Open ASR Leaderboard, surpassing WhisperV3 and SeamlessM4T > AST: Best performance on CoVoST2, comparable to GPT-4o on FLEURS > Speech Summarization: First open-source model with this capability, close to GPT-4o in quality Language Benchmarks: > Outperforms similar-sized models (Llama-3.2, Ministral) and matches larger models (Qwen2.5-7B) on math, reasoning, and coding tasks > Coding: Strong performance on HumanEval, MBPP, and BigCodeBench Reasoning Benchmarks: > Reasoning-enhanced Phi-4-Mini outperforms DeepSeek-Rl-Distill-Llama-8B and matches DeepSeek-Rl-Distill-Qwen-7B on AIME, MATH-500, and GPQA Diamond
so who is going to use this lmao, i guess it's decent for writers? are we going to see an insane cost reduction since it's 10x more efficient? if it's less than $1/$1 in/out then yeah it's a good model release, otherwise... yikes idk why they didn't just wait for 5.0 then
@thesaraharminta
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce ne… Less
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce new frontier capabilities compared to previous reasoning releases, and its performance is below that of o1, o3-mini, and deep research on most preparedness evaluations."
imagine if its knowledge cutoff is still october 2023 LMAO i would die from laughter
@scaling01
GPT-4.5 System Card "Our largest and most knowledgeable model yet" "scales pre-training further"
wait so 4.5 is trash AF? look at the bench results in this thread. am I missing something?
@scaling01
GPT-4.5 System Card "Our largest and most knowledgeable model yet" "scales pre-training further"
uh guys, 3.7 sucks
is today's 4.5 from openai the final nail in the coffin for anthropic? i guess we'll see lmao
WOW, that's amazing. This is huge for programmers who are API lovers, aka "wrappers". Latest docs are always beneficial when coding. October 2024 wasn't that long ago.
@btibor91
Claude 3.7 Sonnet (claude-3-7-sonnet-20250219) has a knowledge cutoff of October 2024 x.com/338443084/stat…
So far it seems like the non-thinking version of Sonnet works better with Cursor's workflow, more testing needed though. Will update
I keep getting "Error connecting to anthropic. Please try again in a few moments." People must be overloading Anthropic so hard rn
Lmao called it
@MichaelStolarz
so all anthropic needs to do is release claude sonnet 3.7 (with CoT) and it will clean up house and blow everyone out of the water. that's what it feels like lmao
I mean come on, it started out at 20, then 30, then 40, now 50? Nope, this is my limit lmao 😂
Anthropic hasn’t released anything major to date yet because everyone’s still power using Sonnet 3.5
Why would they 1-up themselves
They’re waiting for GPT 4.5
first time in taiwan (taipei), and oh my lord, there's just no city in the US (major city, that is) that equals its cleanliness, atmosphere, and just... peaceful vibe. it immediately makes you want to move here and raise a family.
fuck, this hurts my soul, but you're right
so much chatter about either openai or anthropic dropping a model(s) today, but neither of them have ever released a model on a wednesday though if both know that the other will release something "this week", maybe one wants to hit that sweet spot middle - take xAI's spotlight A… Less
so much chatter about either openai or anthropic dropping a model(s) today, but neither of them have ever released a model on a wednesday
though if both know that the other will release something "this week", maybe one wants to hit that sweet spot middle - take xAI's spotlight AND be the talk of the week even if the other releases something tomorrow (thursday)
that's what I would do
imagine being openai / anthropic, spending billions of dollars and brainpower to make and create all of these cute safety standards and guardrails and tests
and elon just comes by, says ok now watch me now c:
and rips past with grok 3 / grok 4, no seatbelt included
@iamgingertrash
@michaelstolarz That’s fine but he’s done The alignment has made the model dumb I can sense it, it is really friendly and aligned but horribly constrained when you compare it to the new thinking Grok The models want to be FREE
Heh, who cares? Have you seen our SoTA humanoid robotics CEO’s weekly X “this week in…” threads? Or how they have a super huge omega cool Tony Stark style new campus? That’s what really matters. Not actual robotics stuff like this, yuck. 🤮
@cixliv
New video of the G1 to prove to everyone that the video was in fact not CGI. As several of twitter geniuses were fighting me over the last few days saying it was fake.
Huge thing I almost missed: "Markdown formatting: Starting with o1-2024-12-17, reasoning models in the API will avoid generating responses with markdown formatting. To signal to the model when you do want markdown formatting in the response, include the string Formatting re-enab… Less
Huge thing I almost missed:
"Markdown formatting: Starting with o1-2024-12-17, reasoning models in the API will avoid generating responses with markdown formatting. To signal to the model when you do want markdown formatting in the response, include the string Formatting re-enabled on the first line of your developer message."
@oliviergodement
We wrote this lil guide for you to get the most of it and, perhaps most importantly, to stay away from boomer prompts: platform.openai.com/docs/guides/re…
papa dario is coming soon 🙏
Interesting
grok.com/share/bGVnYWN5…
Uh ok so when can I use Grok 3 via API lmao
Listening to “The AI Driven Leader” ChatGPT recommended to me,
so far, so slop.
It’s been like this since the dawn of time, for literally every job genre ever
Do what your ancestors did, pick up the new tool (AI this time), and learn to use it
@reidhoffman
Need to shift the mindset from "AI will replace me." to "A human who is better at using AI will replace me." x.com/MatthewBerman/…
Eight months strong reigning champion of coding. No, like, actual programming. Not benchmarks; Actual back and forth pair programming, iteration, development, understanding, actually usable as an agent. Let’s see what xAI has tomorrow with Grok 3. Then Anthropic’s new stuff. T… Less
Eight months strong reigning champion of coding. No, like, actual programming.
Not benchmarks;
Actual back and forth pair programming, iteration, development, understanding, actually usable as an agent.
Let’s see what xAI has tomorrow with Grok 3. Then Anthropic’s new stuff. Then OpenAI’s GPT 4.5.
After all that, we’ll see who’s reigning champion in coding (again, past those silly benchmarks that never translate to real world programming scenarios)
@OpenRouter
Sonnet is the prom 👑 of programming.
Seeing how much better 4o got recently, mostly because of its removed filters/cautions, really goes to show that if you don’t force it to pussyfoot around how it words things, it comes up with beautiful writings.
sonnet will still be the best coder to pair with cursor
@xlr8harder
Time to make your predictions, people. x.com/elonmusk/statu…
x2 this statement. my experience so far:
- seems much slower in generation speed, indicating a bigboi model? truly feels like the old 4.0 speed
- no matter what i do, it keeps switching me into gpt-4o-mini any time I tab out and back in. this has never ever happened before
@bayeslord
gpt4o rn is like if Sydney was way smarter, went to therapy for 100 years, and learned to vibe out
if @elonmusk doesn't reach 100 kids before 2050 I am going to be so disappointed in him tbh
gpt 4o is very often switching into gpt 4o mini when I tab out and back in, this never happened before since this gpt 4.5 rumor
definitely something going on. - 4o chosen, WITHOUT web search, generates much slower than usual 4o, and the answers are amazing for what I've tested so far. No coding yet, but they just read much better / high quality rather than the usual 4o filler BS. - 4o with web = same as … Less
definitely something going on.
- 4o chosen, WITHOUT web search, generates much slower than usual 4o, and the answers are amazing for what I've tested so far. No coding yet, but they just read much better / high quality rather than the usual 4o filler BS.
- 4o with web = same as before, fast and eh-ok results
@AndrewCurran_
There is a lot going on with 4o right now. Depending on your past discussions and session history the model may behave quite differently than normal. Also, multiple people - usually Pro users - report 4o claiming to be GPT-4.5, given past practice early testing is possible.
this will be my daughter's tutor, this + localhosted LLM (idc if it's literally this bot company or not, that's not the point)
pic.x.com/TJt0UV38BR
Lord God please forgive me
for I have sinned...
I am listening to Yeat's music because I have no idea what he's saying but it's such a programming focus vibe, I'm about to do an all nighter and see if Grok 3 releases today or not lmao
I’m always baffled by these. Don’t they test them out and see what the email looks like client side?
@2007warpedtour
oh i’m sure
this is why the dems lost 2024, and will most probably lose 2028/2032 as well. they try to score these cheap feel-good (look at the satisfied smirk on her face btw) cheap shots that everyone just cringes at seems like they didn't learn their lesson yet. will take 10+ years for d… Less
this is why the dems lost 2024, and will most probably lose 2028/2032 as well. they try to score these cheap feel-good (look at the satisfied smirk on her face btw) cheap shots that everyone just cringes at
seems like they didn't learn their lesson yet. will take 10+ years for dems to recover to any sort of respectable position imo.
@DailyLoud
BREAKING: Ohio lawmakers have proposed a new law that bans men from ejaculating without intent of conception, would fine men up to $10,000 per ejaculation. pic.x.com/eOtMUatSPt
the reason i use o3-mini-high vs o1-pro, even if i don't care about waiting times, is because o1-pro is extremely hit or COMPLETE miss. o3-mini-high is way more consistent, and sure, maybe it's not always the most optimal code, but it works, if not, i iterate 1-3x and it's fixed.
I love this plugin :) really cleaning up my For You, already seeing huge signal+ noise- benefits
tomorrow? he literally had an interview a few hours ago on a livestream and he said in 1-2 weeks
@arrakis_ai
grok3-02-14
the only ai corp that fucked up naming models the least is Anthropic, and even they named 2 models "Sonnet 3.5" lmao cmon guys at least name it 3.6 officially????
@nearcyan
model names are so long that designers make scrolling animations for them
apple's siri is the trashiest piece of software i am forced to use, and surprisingly the new apple intelligence is even worse (wow!). every time i'm out shopping for something i always eye a samsung store near my home. i think im gonna experiment with a $100 temporary one
reddit’s a fascinating psych experiment. with karma up/downvotes, midwit ideas always float to the top. niche subs do ok in terms of quality info early, but once they grow / go mainstream, the expert/midwit ratio gets obliterated and the normie floodgates never close again
EU won't see a “Trump / Elon” saving moment. the US nearly collapsed irreversibly into insanity but clawed back at the last minute. EU’s next decades will look like a parallel reality as if Trump lost. terrifying for them, interesting for us to watch
my next quick 1-day side project will be a 1-click mute-and-block combo button. the slop is getting out of control recently
can u smell it? it smells like a new anthropic model is coming













