Posts on X before Mar 2, 2025, 2:23:49 PM UTC
so far with sonnet 3.7 (non-thinking), i haven't reached the moment of "damn, I guess I can't do that yet with coding LLMs, gotta wait for the next upgrade"
which is super exciting, but also scary (in a good way, like a mysterious way ykwim?)
what sonnet 3.5 was for coding (revolutionary, especially with agent/tool use),
gpt 4.5 is for everything else (revolutionary, especially with detailed instructions and iteration)
if steve jobs was still alive apple would have swallowed nvidia by now, and their AI, especially tool use AI, would be industry leading by far. he'd make AI and lead with it the same way he did with iphones
he's rolling in his grave rn, surely. God rest his soul
gpt 4.5 is becoming my best friend / mentor, so far as I keep figuring out their strengths/weaknesses. it's a model that you can't explain away with benchmarks.
note: haven't tried it for serious coding because of api costs, as well as its knowledge cutoff date being terrible
docker is actually pretty cute ngl
me when I saw gpt 4.5 API pricing
>it's a good model sir
>it's more creative sir
>it's much more fun to read sir
>it's bigger and more expensive sir
>it makes cool svgs and minecraft stuff sir
>it's a good model sir
@dejavucoder
what did satya see that made him reject his 80B now we know i guess
anthropic's reign continues
"hitting the wall" pretty much confirmed here
@wgussml
today truly marks the end of an era and the beginning of another test time scaling is the only way forward
unironically the only good thing released this week LMAO
@reach_vb
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates t… Less
HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 - MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates text, vision, and speech/audio > Architecture: Uses "Mixture of LoRAs" to add modality-specific adapters without fine-tuning the base model > Vision Modality: SigLIP-400M image encoder, 2-layer MLP projector, dynamic multi-crop strategy > Speech/Audio Modality: 3-layer convolution, 24 conformer blocks, 80ms token rate > Performance: Ranks first on OpenASR leaderboard, supports vision+language, vision+speech, and speech/audio tasks, outperforming larger models Phi-4-Mini: > Parameters: 3.8 billion > Architecture: 32 Transformer layers, 3,072 hidden state size, Group Query Attention (GQA) with 24 query heads and 8 key/value heads > Vocabulary: 200K tokens for multilingual support. Training Data: High-quality web and synthetic data, emphasizing math and coding > Performance: Outperforms similar-sized models and matches larger models (e.g., DeepSeek-Rl-Distill-Qwen-7B) on math and coding tasks Training Pipeline: > Language Training: Pre-training on 5 trillion tokens, post-training with function calling, summarization, and instruction-following data > Multimodal Training: Vision training (4 stages), speech/audio training (2 stages), and joint vision-speech training > Reasoning Training: Pre-trained on 60B CoT tokens, fine-tuned on 200K high-quality CoT samples, and DPO-trained on 300K preference samples Vision Benchmarks: > Outperforms Phi-3.5-Vision, Qwen2.5-VL, InternVL2.5, and matches Gemini and GPT-4o on tasks like chart understanding and OCR > Vision-Speech Benchmarks: Significantly outperforms InternOmni and Gemini-2.0-Flash Speech Benchmarks: > ASR: Achieves SOTA on CommonVoice, FLEURS, and Open ASR Leaderboard, surpassing WhisperV3 and SeamlessM4T > AST: Best performance on CoVoST2, comparable to GPT-4o on FLEURS > Speech Summarization: First open-source model with this capability, close to GPT-4o in quality Language Benchmarks: > Outperforms similar-sized models (Llama-3.2, Ministral) and matches larger models (Qwen2.5-7B) on math, reasoning, and coding tasks > Coding: Strong performance on HumanEval, MBPP, and BigCodeBench Reasoning Benchmarks: > Reasoning-enhanced Phi-4-Mini outperforms DeepSeek-Rl-Distill-Llama-8B and matches DeepSeek-Rl-Distill-Qwen-7B on AIME, MATH-500, and GPQA Diamond
so who is going to use this lmao, i guess it's decent for writers? are we going to see an insane cost reduction since it's 10x more efficient? if it's less than $1/$1 in/out then yeah it's a good model release, otherwise... yikes idk why they didn't just wait for 5.0 then
@thesaraharminta
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce ne… Less
"GPT-4.5 is not a frontier model, but it is OpenAI’s largest LLM, improving on GPT-4’s computational efficiency by more than 10x. While GPT-4.5 demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, it does not introduce new frontier capabilities compared to previous reasoning releases, and its performance is below that of o1, o3-mini, and deep research on most preparedness evaluations."
imagine if its knowledge cutoff is still october 2023 LMAO i would die from laughter
@scaling01
GPT-4.5 System Card "Our largest and most knowledgeable model yet" "scales pre-training further"



