Posts on X before Dec 26, 2024, 1:23:30 PM UTC
zuck yelling at his team rn about deepseek v3
wow, @warpdotdev's new Next Command feature is a HUGE speedup, it's really accurate, even accounts for all of my aliases and randomness in my everyday workflow
"this is the worst it will ever be" i can't imagine how this feature will blossom in the next updates. 10/10 already
openai search vs grok search
and you're still using something other than grok search?
imagine what grok 3 search is going to be like
cc @teortaxesTex @TheXeophon @gengjiawen @elonmusk
you ever stumble upon a very niche piece of the internet, see unfathomable things going on where people are obsessed about some specific thing, and wonder "how do these people make enough money for them to have so much time to spend arguing about these things?"
this is what o3 is doing rn in case you’re wondering
@tszzl
my strong belief is that you can scale the most clunky piece of shit algorithm to agi and beyond and then ask it how intelligence actually works and realize you built a massive 40’s style ugly vacuum tube mainframe and the asi patiently explains that transistors exist
After almost 10 years of not using Reddit, I decided to create an account there today and see if I’ve been missing out on any alpha in my interests
One word answer: NOPE
I wonder if OpenAI is using o3 to sift through o3/o3-mini safety testing applications 😂
Current programming king to this day. 6 months of uncontested domination.
So now they can use o3 to make o1 as cheap as sonnet 3.6 right? RIGHT?!
This is the scariest tweet I’ve ever read. Good scary, but scary.
@polynoamial
We announced @OpenAI o1 just 3 months ago. Today, we announced o3. We have every reason to believe this trajectory will continue.
GPT 4o = first oh
GPT 4o mini = second oh
GPT 4.5 oh = third oh?
LMAO
Ok so finally a Sonnet 3.6 competitor, yay!
@sama
ho ho ho 🎅 see you tomorrow
OpenAI vs. Google vs. Anthropic
circa ending of 2024
they realized that the audience they were pandering to with limitations/restrictions weren't using their releases to begin with
@skcd42
what changed with Google. They are literally killing it with the releases
"judge me not by my youtube recommended slop, but by my x for you page"
-me
holy fuck @XDevelopers @elonmusk COOKED with this new For You algo update. it's amazing now, best algo ever, so much alpha on my timeline now from randos (that I now follow!)
i'm in a mental dilemma;
- if openai doesn't release gpt 4.5 or whatever (their response to claude's sonnet 3.6 current dominance), i'm bearish on them
BUT ALSO
- if anthropic doesn't release something to outperform o1/an even better sonnet 3.6/opus 3.5, then im bearish too...
My wife just called LinkedIn, "link it in". And from now on, that's how I'll pronounce it too.
sometimes when I use gpt-4o by mistake, the response feels worse than what I'd have expected back in the gpt-3.5-turbo days
We still make beautiful things, they’re just so small you need an electron microscope to appreciate their structural beauty
@thegenesisbl0ck
We used to create such beautiful things, now we make rap songs about drugs and pu$$y.
yawn... another day where grok @xai completely mogs chatgpt search
The GOAT does it again.
@kepano
Thanks for reaching out. Do you mind if I share this email on social media to gauge the reaction from our community?
Still the case to this day
I gauge this by which model I use for writing and coding in Cursor (without doing the usage-based pricing BS regarding trying to use o1 models in Cursor)
So yeah sonnet is still king
@Teknium
Cant believe OAI has let sonnet 3.5 be the best (not inference scaled, but even that in many areas) model on the market for so long
2025 is the year Google and @OfficialLoganK brutally mog everyone else
@RubenEVillegas
A cat roars while looking at its reflection in the mirror but instead sees itself as a lion roaring #veo2
the winner of the "native multimodality" wars is going to be the first model that can tap in to the latest web results
from personal testing the leaders in this space are google obviously, and SURPRISINGLY, grok
chatgpt's search is 3rd place for me rn
ok im consistently hitting 7 minute thinking times with o1 pro, that means we're REALLY cooking, the results are immaculate tbh
guys o1 pro has been thinking for 6 mins now what's it concocting for me im scared
ChatGPT Search vs xAI's Grok search. The results speak for themselves. Holy fuck what a difference in accuracy and conciseness.
every morning I wake up and am sad again after finding out there are no new online models via API other than the old Perplexity 3.1 / Sonar stuff
such an easy W for whoever releases this. though I must say, OpenAI is quite primed for this, just make ChatGPT's "Search" API-ready
uh... arc? baby? 🥲
llama-3.3-70b is an absolute beast. not enough people talking about it imo. so accurate, follows guidelines and rules so well, rarely defects from its expected outputs. amazing.
if you are prototyping a business idea and it's taking you more than one week to have a good rough version to test out with yourself and your friends, then you are not up to speed with how fast development is these days
my first sora generation... noice.
ok and now it's a love-only relationship after I got near instant support in the DMs and got my developer profile turned on by them
absolutely amazing experience so far. you can just do things
@MichaelStolarz
So I have a love and hate relationship with @GroqInc - long story short, I wish their developer tier finally came out so I can buy credits or whatever so that it's finally reliable rather than being randomly limited since there's only a free tier.
If you're using OpenRouter, make sure that you are setting quantizations to BF16 and then down from there so that your requests are always hitting the best available version, even if it may be a little bit slower.
The results difference speaks for itself.
So I have a love and hate relationship with @GroqInc - long story short, I wish their developer tier finally came out so I can buy credits or whatever so that it's finally reliable rather than being randomly limited since there's only a free tier.
We are so back. Holy fuck 2025 is going to be insane.
Google finally got their shit together it seems; this is their FLASH model, its multimodal, speech output, video intake, cheap AF.
Google is posturing to be not only #1, but #1 BY A LARGE MARGIN.
@OfficialLoganK
Gemini 2.0 Flash is our strongest model to date, crazy progress for a small model like Flash : )
This is already solved in the labs, they’re just packing it up for consumers rn.
We’ll see the first usable ones late Q1/Q2 which will still be clunky and buggy; then late 2025 it should be as “game changing” as the first useful LLMs felt
@GregKamradt
LLMs don't have long term memory - they're stateless Who's doing memory as a service? What's the best practice with incremental data and adding it to a long term memory?
Advanced Voice feels like I’m just talking with GPT 3.5 Turbo. Nothing of value is gained talking to it, it’s just a feel good loop that FEELS useful but it’s actually just a waste of time
@SmokeAwayyy
ChatGPT Advanced Voice Impressions After ~10 hours of 4o voice chats using ChatGPT Pro here are some initial thoughts: - Pretty good. Not great. - Responds too quickly. Turns out responding instantly isn't actually desirable. It fails to detect contextual pauses mid sentence/th… Less
ChatGPT Advanced Voice Impressions After ~10 hours of 4o voice chats using ChatGPT Pro here are some initial thoughts: - Pretty good. Not great. - Responds too quickly. Turns out responding instantly isn't actually desirable. It fails to detect contextual pauses mid sentence/thought so you need to have your entire thought formed before speaking. I would rather speak with an o1 type model that at least takes a second or two to reason. - Needs an optional toggle to speak up during long pauses. Currently it just sits there forever if you don't say anything. - Emotionless. I know "AI is a tool" and all that, but 4o feels completely emotionless and too rigid. - Way too strict. OpenAI models are now a minefield of compliance. 4o Voice is the most strict (won't output). 4o Text is less strict (warning label but outputs). o1 Text is now the least strict model because it is able to think around the OpenAI policies and self-moderate at a human level (outputs with less warning labels). - Incredible at speaking in different languages, accents, and styles. - Can't hold unique voices consistently. It constantly needs to be reminded to maintain a new voice/language. - Responses are all very short. Hard to get it to say more than a paragraph in one response. - Speaks much slower/differently at the end of long chats. - Chats are limited to 1 hour each so 24/7 voice mode is currently not possible even with ChatGPT Pro. Also have a 4 limit placed on advanced voice during a normal mid day chat so I'm guessing it's not actually unlimited. The ideal Voice Mode in my mind would look something like unlimited o1 with reasoning, custom instructions, "infinite" memory, search, and vision.
Insane how we still don’t have an online Search LLM API using the SotA models. @perplexity_ai @OpenAI please 🙏
Holy fuck… I don’t think people understand what this means if it’s true. I’ve used this experimental model a lot and I thought it’s just another Gemini pro 1.5 upgrade they’re pushing out before 2.0. But nope, seems like it’s possibly FLASH 2.0. Basically free imo. … Less
Holy fuck… I don’t think people understand what this means if it’s true. I’ve used this experimental model a lot and I thought it’s just another Gemini pro 1.5 upgrade they’re pushing out before 2.0.
But nope, seems like it’s possibly FLASH 2.0. Basically free imo. x.com/legit_rumors/s…
for the past 2 days i've been like "damn the For You algo has gotten really, really good" but today I realized I've just been in the Following tab, lmao
surely Anthropic releases a new model / another sonnet update during these 12 days of OpenAI Christmas?
o1 pro just isn't that useful unless you're doing college level mathematics / science based reasoning.
it's not better for code, and the tradeoff in speed (~3-5 seconds with o1 vs 2-3 minutes with o1-pro) isn't worth it, especially if you need to do iterations/refactors.
in less than 12 hours we're going to see what @sama is so excited about
o1 pro shouldn't be your programmer, it should be your architect
yeah I mean $200 per month for pro is worth it even if the only benefit was unlimited o1.
even if there wasn't o1 pro, or unlimited voice mode - just having unlimited o1 is a game changer imo
one thing i've noticed so far, using nothing but o1-pro for an entire day, and then now using o1 for an entire day
is that iteration changes and your workflow in general is much faster with o1
I think I'm going to use o1 pro only for very complicated planning/decision trees
it's so interesting to see that sometimes when i try to do something in o1-pro, it just can't get it right but then i pop it into o1, and it's perfect on the first try i wonder what that means... is o1-pro just overthinking everything? gets stuck on something? maybe its contex… Less
it's so interesting to see that sometimes when i try to do something in o1-pro, it just can't get it right
but then i pop it into o1, and it's perfect on the first try
i wonder what that means... is o1-pro just overthinking everything? gets stuck on something?
maybe its context window with all of its behind-the-scenes think-through stuff gets all too big, and it just forgets what it's supposed to do?
very curious to keep testing this out 🤔
i have very mixed feelings regarding O1 Pro. on one hand, when it gets it right, it gets it right beyond what you expected. as in, it does things that you didn't even ask for, which is great for certain situations, but terrible and time-wasting for others, because then you have … Less
i have very mixed feelings regarding O1 Pro.
on one hand, when it gets it right, it gets it right beyond what you expected. as in, it does things that you didn't even ask for, which is great for certain situations, but terrible and time-wasting for others, because then you have to go back and modify the prompt and specify, hey, don't go ahead and do this, or don't change this.
so suffice to say, it's hit or miss, except when you hit, it's like you hit the jackpot. and when you miss, it's just a mild annoyance.
is it worth $200 a month? in my opinion, a quick good gauge for yes/no is if you are able to afford food delivery every day, if you're at that level of disposable income, then yes, you must find a way to rebudget a few things and get this.
of course, this should go without saying, but if you do end up buying the subscription, you should be using it as many times as you can every single day.
i'm slowly getting to the point where instead of browsing twitter, i'm browsing O1 Pro. or more specifically, i prompt O1 Pro and then open up my other tab into just typical O1, which responds much faster now. and then i chat there while i wait for O1 Pro, and then back and forth and back and forth, you get the idea.






















