Posts on X before Dec 17, 2024, 12:00:28 PM UTC

the winner of the "native multimodality" wars is going to be the first model that can tap in to the latest web results from personal testing the leaders in this space are google obviously, and SURPRISINGLY, grok chatgpt's search is 3rd place for me rn
ok im consistently hitting 7 minute thinking times with o1 pro, that means we're REALLY cooking, the results are immaculate tbh
ChatGPT Search vs xAI's Grok search. The results speak for themselves. Holy fuck what a difference in accuracy and conciseness.
Media attached to this postMedia attached to this post
every morning I wake up and am sad again after finding out there are no new online models via API other than the old Perplexity 3.1 / Sonar stuff such an easy W for whoever releases this. though I must say, OpenAI is quite primed for this, just make ChatGPT's "Search" API-ready
llama-3.3-70b is an absolute beast. not enough people talking about it imo. so accurate, follows guidelines and rules so well, rarely defects from its expected outputs. amazing.
if you are prototyping a business idea and it's taking you more than one week to have a good rough version to test out with yourself and your friends, then you are not up to speed with how fast development is these days
ok and now it's a love-only relationship after I got near instant support in the DMs and got my developer profile turned on by them absolutely amazing experience so far. you can just do things
Media attached to this postMedia attached to this post
So I have a love and hate relationship with @GroqInc - long story short, I wish their developer tier finally came out so I can buy credits or whatever so that it's finally reliable rather than being randomly limited since there's only a free tier.
If you're using OpenRouter, make sure that you are setting quantizations to BF16 and then down from there so that your requests are always hitting the best available version, even if it may be a little bit slower. The results difference speaks for itself.
Media attached to this post
So I have a love and hate relationship with @GroqInc - long story short, I wish their developer tier finally came out so I can buy credits or whatever so that it's finally reliable rather than being randomly limited since there's only a free tier.
We are so back. Holy fuck 2025 is going to be insane. Google finally got their shit together it seems; this is their FLASH model, its multimodal, speech output, video intake, cheap AF. Google is posturing to be not only #1, but #1 BY A LARGE MARGIN.
Gemini 2.0 Flash is our strongest model to date, crazy progress for a small model like Flash : )
Media attached to this post
This is already solved in the labs, they’re just packing it up for consumers rn. We’ll see the first usable ones late Q1/Q2 which will still be clunky and buggy; then late 2025 it should be as “game changing” as the first useful LLMs felt
LLMs don't have long term memory - they're stateless Who's doing memory as a service? What's the best practice with incremental data and adding it to a long term memory?
Advanced Voice feels like I’m just talking with GPT 3.5 Turbo. Nothing of value is gained talking to it, it’s just a feel good loop that FEELS useful but it’s actually just a waste of time
ChatGPT Advanced Voice Impressions After ~10 hours of 4o voice chats using ChatGPT Pro here are some initial thoughts: - Pretty good. Not great. - Responds too quickly. Turns out responding instantly isn't actually desirable. It fails to detect contextual pauses mid sentence/th… MoreLess
ChatGPT Advanced Voice Impressions After ~10 hours of 4o voice chats using ChatGPT Pro here are some initial thoughts: - Pretty good. Not great. - Responds too quickly. Turns out responding instantly isn't actually desirable. It fails to detect contextual pauses mid sentence/thought so you need to have your entire thought formed before speaking. I would rather speak with an o1 type model that at least takes a second or two to reason. - Needs an optional toggle to speak up during long pauses. Currently it just sits there forever if you don't say anything. - Emotionless. I know "AI is a tool" and all that, but 4o feels completely emotionless and too rigid. - Way too strict. OpenAI models are now a minefield of compliance. 4o Voice is the most strict (won't output). 4o Text is less strict (warning label but outputs). o1 Text is now the least strict model because it is able to think around the OpenAI policies and self-moderate at a human level (outputs with less warning labels). - Incredible at speaking in different languages, accents, and styles. - Can't hold unique voices consistently. It constantly needs to be reminded to maintain a new voice/language. - Responses are all very short. Hard to get it to say more than a paragraph in one response. - Speaks much slower/differently at the end of long chats. - Chats are limited to 1 hour each so 24/7 voice mode is currently not possible even with ChatGPT Pro. Also have a 4 limit placed on advanced voice during a normal mid day chat so I'm guessing it's not actually unlimited. The ideal Voice Mode in my mind would look something like unlimited o1 with reasoning, custom instructions, "infinite" memory, search, and vision.
Holy fuck… I don’t think people understand what this means if it’s true. I’ve used this experimental model a lot and I thought it’s just another Gemini pro 1.5 upgrade they’re pushing out before 2.0. But nope, seems like it’s possibly FLASH 2.0. Basically free imo. … MoreLess
Holy fuck… I don’t think people understand what this means if it’s true. I’ve used this experimental model a lot and I thought it’s just another Gemini pro 1.5 upgrade they’re pushing out before 2.0. But nope, seems like it’s possibly FLASH 2.0. Basically free imo. x.com/legit_rumors/s…
for the past 2 days i've been like "damn the For You algo has gotten really, really good" but today I realized I've just been in the Following tab, lmao
o1 pro just isn't that useful unless you're doing college level mathematics / science based reasoning. it's not better for code, and the tradeoff in speed (~3-5 seconds with o1 vs 2-3 minutes with o1-pro) isn't worth it, especially if you need to do iterations/refactors.
yeah I mean $200 per month for pro is worth it even if the only benefit was unlimited o1. even if there wasn't o1 pro, or unlimited voice mode - just having unlimited o1 is a game changer imo
one thing i've noticed so far, using nothing but o1-pro for an entire day, and then now using o1 for an entire day is that iteration changes and your workflow in general is much faster with o1 I think I'm going to use o1 pro only for very complicated planning/decision trees
it's so interesting to see that sometimes when i try to do something in o1-pro, it just can't get it right but then i pop it into o1, and it's perfect on the first try i wonder what that means... is o1-pro just overthinking everything? gets stuck on something? maybe its contex… MoreLess
it's so interesting to see that sometimes when i try to do something in o1-pro, it just can't get it right but then i pop it into o1, and it's perfect on the first try i wonder what that means... is o1-pro just overthinking everything? gets stuck on something? maybe its context window with all of its behind-the-scenes think-through stuff gets all too big, and it just forgets what it's supposed to do? very curious to keep testing this out 🤔
i have very mixed feelings regarding O1 Pro. on one hand, when it gets it right, it gets it right beyond what you expected. as in, it does things that you didn't even ask for, which is great for certain situations, but terrible and time-wasting for others, because then you have … MoreLess
i have very mixed feelings regarding O1 Pro. on one hand, when it gets it right, it gets it right beyond what you expected. as in, it does things that you didn't even ask for, which is great for certain situations, but terrible and time-wasting for others, because then you have to go back and modify the prompt and specify, hey, don't go ahead and do this, or don't change this. so suffice to say, it's hit or miss, except when you hit, it's like you hit the jackpot. and when you miss, it's just a mild annoyance. is it worth $200 a month? in my opinion, a quick good gauge for yes/no is if you are able to afford food delivery every day, if you're at that level of disposable income, then yes, you must find a way to rebudget a few things and get this. of course, this should go without saying, but if you do end up buying the subscription, you should be using it as many times as you can every single day. i'm slowly getting to the point where instead of browsing twitter, i'm browsing O1 Pro. or more specifically, i prompt O1 Pro and then open up my other tab into just typical O1, which responds much faster now. and then i chat there while i wait for O1 Pro, and then back and forth and back and forth, you get the idea.
Alright, there goes $200 buckaroos, I'll be using nothing but o1 Pro through the entire next 12ish hours, lets see when/if I hit a rate limit/notification 🫡
I'd be down for a $1,000 / month unlimited API access to o1 pro - (obviously agreeing to their caveat of not using this particular inf. API access for commercial purposes, only self-use/development, duh)
One thing I’ve come to realize (fortunately and unfortunately) is that we’re at a point where we can build anything software wise. And I mean that quite literally. If you know how to use agents, self improvement, and just replace your main driver model with whatever’s the latest … MoreLess
One thing I’ve come to realize (fortunately and unfortunately) is that we’re at a point where we can build anything software wise. And I mean that quite literally. If you know how to use agents, self improvement, and just replace your main driver model with whatever’s the latest SOTA… then yeah, there’s nothing you can’t do It’s still difficult, but for those that weren’t experts in certain domains, it would take years of study to apply things in a new domain. Now that same level of quality work can be done in a few months
Interesting to see that no one has reached limits yet with o1 pro… super interesting. Might just have to try it out
The real white pill is realizing Safari is the best browser if you’re deep into the Apple ecosystem already Few.
Hahahaha nope, THAT mode is for gooberment / internal use only
@sama Does the $200 mode remove the thing where it lectures me on morality instead of answering a fucking question?
personally i don't care much about chatgpt itself, i care about what i'll have access to regarding the API fingers crossed that these new "o1-pro" and "gpt-4.5" will be available through API usage - idc at what price, I just want to test them in my own environment/sys prompts
it really feels like major softwares of our time with 100+ team members are actually just churning out unoptimized "uniquely designed" slop just so that they ship something reminds me of game studios
Is it just me or was 18.1.1 a total turd upgrade. Phone doesn’t work anymore.
this +: - AR (ex. apple vision) - speech to action - 3d scene realtime rendering - suno-level audio generation -character consistency/“models” then all you need is a dedicated empty room to build in; walk around, design a movie, make a game, i mean you’re literally tony stark
announcing Krea Editor. our new editing tool feels like magic. who wants beta access? 👇
i've been seeing a lot of rumblings from openai employees (i analyze their post frequency via their notis being turned on for me) it's matching the same frequencies as whenever there's a release coming, so i'm pretty sure something releases this week, maybe even today (tuesday)
Media attached to this post
TIL perplexity API is pretty useless because there's no Perplexity Pro equivalent (ex. reasoning, ability to use Sonnet 3.6 as the search model, etc) Just some outdated llama 3.1 perplexity online models. booooo 😭 @perplexity_ai @AravSrinivas pls implement, I'll pay w/e
bro, just one more all-nighter, bro. the api rewrite will be so worth it, bro. cleaner codebase, bro. just one more, bro. i swear, bro
surprisingly, the only 99.9% accurate speech-to-text, of course, after i do a post-transcription on it, is still OpenAI's Whisper API. all of these other providers that offer Turbo, like Groq, or any of the versions that are hostable on your own, if it's not that good ol' Whispe… MoreLess
surprisingly, the only 99.9% accurate speech-to-text, of course, after i do a post-transcription on it, is still OpenAI's Whisper API. all of these other providers that offer Turbo, like Groq, or any of the versions that are hostable on your own, if it's not that good ol' Whisper, it's just not as accurate. what i find that happens very often is, it'll just not transcribe a portion of your speech. i'm not sure why this happens. it'll just skip over things, or it'll do the notorious "..." instead of actually transcribing out the words. it's almost as if it's too lazy to do the full transcription sometimes. really interesting behavior. hopefully it's fixed in V4, or whatever. but for now, if i want to have a good focus flow, without having to go back and try and fix transcriptions and such, or even worse, having to re-say everything that i initially said, i'm just going to stick with OpenAI's Whisper API for now. and then of course, Sonnet 3.6 for post-transcription touch-ups.
when i first started using Cursor a little over a year ago, i really felt the "10x programmer" meme. especially since i was already so comfortable with the terminal and setting up automations in a pseudo-agent type of way with file watchers and procedures and loops... you get th… MoreLess
when i first started using Cursor a little over a year ago, i really felt the "10x programmer" meme. especially since i was already so comfortable with the terminal and setting up automations in a pseudo-agent type of way with file watchers and procedures and loops... you get the idea. now, with the latest Cursor update, where it has the agent feature, that initial 10x feels like another 10x. it's a steep learning curve, which isn't obvious from a first glance, even for those that use Cursor or code-gen tools every day. i think 99% of programmers don't utilize Cursor efficiently. and then 99% of the 1% that do utilize Cursor don't really think through how this new agent feature is as big of a game changer as code generation AI is. once people catch on and realize how powerful this is, i don't see any software limitations in terms of what can be created, even if tech progress magically stopped and we never have a better model than Sonnet 3.6 (lmao but you get the point) what an amazing time to be alive, waiting for the next state-of-the-art LLM to further this productivity explosion, and more importantly imo, even better workflows that no one has even thought of yet (but AI will for you, especially when personalization becomes an integral part, not even mentioning continuous learning...)
"Almost every important piece of software from the last 10 years uses Go"
GIF attached to this post
Go is the programming language of the internet and the modern era. Almost every important piece of software from the last 10 years uses Go: - Docker - Kubernetes - Terraform - Prometheus - Grafana - Lets Encrypt - CockroachDB - etcd - Ethereum - Vault - Packer - Caddy
2024 isn't over yet, so i really hope i'm wrong with this next statement, but here it is: the only notable advancements that you can really feel in terms of output improvement have been Sonnet 3.5 and the o1/r1 "reasoning" models. everything else is cool, but tbh it just kind o… MoreLess
2024 isn't over yet, so i really hope i'm wrong with this next statement, but here it is: the only notable advancements that you can really feel in terms of output improvement have been Sonnet 3.5 and the o1/r1 "reasoning" models. everything else is cool, but tbh it just kind of seems slopped together