Posts on X before Dec 21, 2024, 9:06:23 AM UTC

This is the scariest tweet I’ve ever read. Good scary, but scary.
We announced @OpenAI o1 just 3 months ago. Today, we announced o3. We have every reason to believe this trajectory will continue.
Media attached to this postMedia attached to this post
they realized that the audience they were pandering to with limitations/restrictions weren't using their releases to begin with
what changed with Google. They are literally killing it with the releases
i'm in a mental dilemma; - if openai doesn't release gpt 4.5 or whatever (their response to claude's sonnet 3.6 current dominance), i'm bearish on them BUT ALSO - if anthropic doesn't release something to outperform o1/an even better sonnet 3.6/opus 3.5, then im bearish too...
sometimes when I use gpt-4o by mistake, the response feels worse than what I'd have expected back in the gpt-3.5-turbo days
We still make beautiful things, they’re just so small you need an electron microscope to appreciate their structural beauty
We used to create such beautiful things, now we make rap songs about drugs and pu$$y.
Media attached to this postMedia attached to this postMedia attached to this post
The GOAT does it again.
Thanks for reaching out. Do you mind if I share this email on social media to gauge the reaction from our community?
From: Bending Spoons
Subject: This might make you really suspicious
Date: Tue, Dec 17, 2024 at 1:20 PM

I am Alexandra from Bending Spoons, the biggest European app developer and publisher.

I know what I am going to say may sound crazy and catch you off guard. How ridiculous would it be to discuss Bending Spoons potentially acquiring Obsidian?

Obsidian actually listens to its users, which has fostered a tight-knit vibrant community. We like products with strong fundamentals and would love to explore the possibility of being part of Obsidian's journey. Do you hate the idea? Let me know if you are available for a short call to discuss more about it.

To be clear, we're for the long term. We'd truly appreciate being on your radar and building a connection.
Still the case to this day I gauge this by which model I use for writing and coding in Cursor (without doing the usage-based pricing BS regarding trying to use o1 models in Cursor) So yeah sonnet is still king
Cant believe OAI has let sonnet 3.5 be the best (not inference scaled, but even that in many areas) model on the market for so long
the winner of the "native multimodality" wars is going to be the first model that can tap in to the latest web results from personal testing the leaders in this space are google obviously, and SURPRISINGLY, grok chatgpt's search is 3rd place for me rn
ok im consistently hitting 7 minute thinking times with o1 pro, that means we're REALLY cooking, the results are immaculate tbh
ChatGPT Search vs xAI's Grok search. The results speak for themselves. Holy fuck what a difference in accuracy and conciseness.
Media attached to this postMedia attached to this post
every morning I wake up and am sad again after finding out there are no new online models via API other than the old Perplexity 3.1 / Sonar stuff such an easy W for whoever releases this. though I must say, OpenAI is quite primed for this, just make ChatGPT's "Search" API-ready
llama-3.3-70b is an absolute beast. not enough people talking about it imo. so accurate, follows guidelines and rules so well, rarely defects from its expected outputs. amazing.
if you are prototyping a business idea and it's taking you more than one week to have a good rough version to test out with yourself and your friends, then you are not up to speed with how fast development is these days
ok and now it's a love-only relationship after I got near instant support in the DMs and got my developer profile turned on by them absolutely amazing experience so far. you can just do things
Media attached to this postMedia attached to this post
So I have a love and hate relationship with @GroqInc - long story short, I wish their developer tier finally came out so I can buy credits or whatever so that it's finally reliable rather than being randomly limited since there's only a free tier.
If you're using OpenRouter, make sure that you are setting quantizations to BF16 and then down from there so that your requests are always hitting the best available version, even if it may be a little bit slower. The results difference speaks for itself.
Media attached to this post
So I have a love and hate relationship with @GroqInc - long story short, I wish their developer tier finally came out so I can buy credits or whatever so that it's finally reliable rather than being randomly limited since there's only a free tier.
We are so back. Holy fuck 2025 is going to be insane. Google finally got their shit together it seems; this is their FLASH model, its multimodal, speech output, video intake, cheap AF. Google is posturing to be not only #1, but #1 BY A LARGE MARGIN.
Gemini 2.0 Flash is our strongest model to date, crazy progress for a small model like Flash : )
Media attached to this post
This is already solved in the labs, they’re just packing it up for consumers rn. We’ll see the first usable ones late Q1/Q2 which will still be clunky and buggy; then late 2025 it should be as “game changing” as the first useful LLMs felt
LLMs don't have long term memory - they're stateless Who's doing memory as a service? What's the best practice with incremental data and adding it to a long term memory?
Advanced Voice feels like I’m just talking with GPT 3.5 Turbo. Nothing of value is gained talking to it, it’s just a feel good loop that FEELS useful but it’s actually just a waste of time
ChatGPT Advanced Voice Impressions After ~10 hours of 4o voice chats using ChatGPT Pro here are some initial thoughts: - Pretty good. Not great. - Responds too quickly. Turns out responding instantly isn't actually desirable. It fails to detect contextual pauses mid sentence/th… MoreLess
ChatGPT Advanced Voice Impressions After ~10 hours of 4o voice chats using ChatGPT Pro here are some initial thoughts: - Pretty good. Not great. - Responds too quickly. Turns out responding instantly isn't actually desirable. It fails to detect contextual pauses mid sentence/thought so you need to have your entire thought formed before speaking. I would rather speak with an o1 type model that at least takes a second or two to reason. - Needs an optional toggle to speak up during long pauses. Currently it just sits there forever if you don't say anything. - Emotionless. I know "AI is a tool" and all that, but 4o feels completely emotionless and too rigid. - Way too strict. OpenAI models are now a minefield of compliance. 4o Voice is the most strict (won't output). 4o Text is less strict (warning label but outputs). o1 Text is now the least strict model because it is able to think around the OpenAI policies and self-moderate at a human level (outputs with less warning labels). - Incredible at speaking in different languages, accents, and styles. - Can't hold unique voices consistently. It constantly needs to be reminded to maintain a new voice/language. - Responses are all very short. Hard to get it to say more than a paragraph in one response. - Speaks much slower/differently at the end of long chats. - Chats are limited to 1 hour each so 24/7 voice mode is currently not possible even with ChatGPT Pro. Also have a 4 limit placed on advanced voice during a normal mid day chat so I'm guessing it's not actually unlimited. The ideal Voice Mode in my mind would look something like unlimited o1 with reasoning, custom instructions, "infinite" memory, search, and vision.
Holy fuck… I don’t think people understand what this means if it’s true. I’ve used this experimental model a lot and I thought it’s just another Gemini pro 1.5 upgrade they’re pushing out before 2.0. But nope, seems like it’s possibly FLASH 2.0. Basically free imo. … MoreLess
Holy fuck… I don’t think people understand what this means if it’s true. I’ve used this experimental model a lot and I thought it’s just another Gemini pro 1.5 upgrade they’re pushing out before 2.0. But nope, seems like it’s possibly FLASH 2.0. Basically free imo. x.com/legit_rumors/s…
for the past 2 days i've been like "damn the For You algo has gotten really, really good" but today I realized I've just been in the Following tab, lmao
o1 pro just isn't that useful unless you're doing college level mathematics / science based reasoning. it's not better for code, and the tradeoff in speed (~3-5 seconds with o1 vs 2-3 minutes with o1-pro) isn't worth it, especially if you need to do iterations/refactors.
yeah I mean $200 per month for pro is worth it even if the only benefit was unlimited o1. even if there wasn't o1 pro, or unlimited voice mode - just having unlimited o1 is a game changer imo
one thing i've noticed so far, using nothing but o1-pro for an entire day, and then now using o1 for an entire day is that iteration changes and your workflow in general is much faster with o1 I think I'm going to use o1 pro only for very complicated planning/decision trees
it's so interesting to see that sometimes when i try to do something in o1-pro, it just can't get it right but then i pop it into o1, and it's perfect on the first try i wonder what that means... is o1-pro just overthinking everything? gets stuck on something? maybe its contex… MoreLess
it's so interesting to see that sometimes when i try to do something in o1-pro, it just can't get it right but then i pop it into o1, and it's perfect on the first try i wonder what that means... is o1-pro just overthinking everything? gets stuck on something? maybe its context window with all of its behind-the-scenes think-through stuff gets all too big, and it just forgets what it's supposed to do? very curious to keep testing this out 🤔
i have very mixed feelings regarding O1 Pro. on one hand, when it gets it right, it gets it right beyond what you expected. as in, it does things that you didn't even ask for, which is great for certain situations, but terrible and time-wasting for others, because then you have … MoreLess
i have very mixed feelings regarding O1 Pro. on one hand, when it gets it right, it gets it right beyond what you expected. as in, it does things that you didn't even ask for, which is great for certain situations, but terrible and time-wasting for others, because then you have to go back and modify the prompt and specify, hey, don't go ahead and do this, or don't change this. so suffice to say, it's hit or miss, except when you hit, it's like you hit the jackpot. and when you miss, it's just a mild annoyance. is it worth $200 a month? in my opinion, a quick good gauge for yes/no is if you are able to afford food delivery every day, if you're at that level of disposable income, then yes, you must find a way to rebudget a few things and get this. of course, this should go without saying, but if you do end up buying the subscription, you should be using it as many times as you can every single day. i'm slowly getting to the point where instead of browsing twitter, i'm browsing O1 Pro. or more specifically, i prompt O1 Pro and then open up my other tab into just typical O1, which responds much faster now. and then i chat there while i wait for O1 Pro, and then back and forth and back and forth, you get the idea.
Alright, there goes $200 buckaroos, I'll be using nothing but o1 Pro through the entire next 12ish hours, lets see when/if I hit a rate limit/notification 🫡
I'd be down for a $1,000 / month unlimited API access to o1 pro - (obviously agreeing to their caveat of not using this particular inf. API access for commercial purposes, only self-use/development, duh)
One thing I’ve come to realize (fortunately and unfortunately) is that we’re at a point where we can build anything software wise. And I mean that quite literally. If you know how to use agents, self improvement, and just replace your main driver model with whatever’s the latest … MoreLess
One thing I’ve come to realize (fortunately and unfortunately) is that we’re at a point where we can build anything software wise. And I mean that quite literally. If you know how to use agents, self improvement, and just replace your main driver model with whatever’s the latest SOTA… then yeah, there’s nothing you can’t do It’s still difficult, but for those that weren’t experts in certain domains, it would take years of study to apply things in a new domain. Now that same level of quality work can be done in a few months
Interesting to see that no one has reached limits yet with o1 pro… super interesting. Might just have to try it out
The real white pill is realizing Safari is the best browser if you’re deep into the Apple ecosystem already Few.
Hahahaha nope, THAT mode is for gooberment / internal use only
@sama Does the $200 mode remove the thing where it lectures me on morality instead of answering a fucking question?
personally i don't care much about chatgpt itself, i care about what i'll have access to regarding the API fingers crossed that these new "o1-pro" and "gpt-4.5" will be available through API usage - idc at what price, I just want to test them in my own environment/sys prompts