Posts on X before Oct 31, 2024, 5:29:42 AM UTC
ok, how about - get this - christianity?
@tsarnick
Bill Gates says AI will become so good at solving problems and creating attractive activities for humans to do that we need a new religion or philosophy to stay connected to one another
new net worth goal: $1 decillion
@CultureCrave
Russia is trying to fine Google $20 decillion over YouTube bans • The fine surpasses the entire wealth and asset value on Earth • Google so far has ignored their demands
insane to me that @cursor_ai STILL doesn't have a commit message generator button like copilot does. is it patented or something?! seems like a no-brainer feature, and also seems like cursor deliberately isn't adding it in for some reason.
probably because python is the best performing language for LLM coding success
@ashtom
Even in this, Python is still growing faster than both JS and TS combined! So there’s cause for celebration from the Pythonistas 🐍🐍🐍 x.com/BenLesh/status…
for the past year or so, 99% of the time the answer to any question like "I'm having a hard time understanding ___, what should I do?" is: Ask a SotA LLM. The latest Sonnet 3.5 from Anthropic, 4o or o1 from OpenAI. They'll explain + create example problems for you. Every single… Less
for the past year or so, 99% of the time the answer to any question like "I'm having a hard time understanding ___, what should I do?" is:
Ask a SotA LLM. The latest Sonnet 3.5 from Anthropic, 4o or o1 from OpenAI.
They'll explain + create example problems for you. Every single time.
if you use LLMs primarily for coding purposes, you're wasting your time looking at any other benchmark other than Aider's. let me explain and show you the best to track. first, a friendly reminder that chatbot arena scores are, for the most part, retarded, most likely skewed wi… Less
if you use LLMs primarily for coding purposes, you're wasting your time looking at any other benchmark other than Aider's.
let me explain and show you the best to track.
first, a friendly reminder that chatbot arena scores are, for the most part, retarded, most likely skewed with sophisticated bot systems to increase votes (and if not, then this is a missed opportunity for all the big labs, but i digress-)
most most importantly - llm leaderboards like lmsys/lmarena/chatbot arena are based on vibes and how cool the output looks (formatting, speech style) instead of "correct" results
the only benchmark i've found to be consistently reliable in how i experience the quality of today's SotA LLMs is the Aider code editing benchmark from @paulgauthier, as well as his second refactoring-based coding benchmark (which is way more demanding).
this is exactly the order of LLMs I would use if the best one were unreachable for some reason. No sonnet? o1-preview then. no o1-preview either? OK, gpt-4o is fine too (opus if the price was the same). and so on. the list is just... perfect.
@arena
Chatbot Arena Update🔥 @AnthropicAI latest Claude 3.5 Sonnet, has been extensively tested in Arena, securing an impressive #6 overall and #3 under style-control! With over 7K community votes, the new Sonnet is showing exceptional strength across various domains. Highlights: - H… Less
Chatbot Arena Update🔥 @AnthropicAI latest Claude 3.5 Sonnet, has been extensively tested in Arena, securing an impressive #6 overall and #3 under style-control! With over 7K community votes, the new Sonnet is showing exceptional strength across various domains. Highlights: - Hard Prompts: #4 (#1 with SC) - Coding: #2 (#1 with SC) - Math: #3 Congrats to @AnthropicAI on the impressive new release! More analysis below👇
new reason to vote for trump: if kamala wins, geohot isn't coming back to america. GEOHOT. it's like if white zombie lost rob zombie.
@geohotarchive
America’s Future geohot.github.io//blog/jekyll/u…
this year was @rlgrime's last halloween album release (13/XIII); time to bring it back all the way to the beginning. let's listen to and rate all of them from the start. I'll rate each on a scale of 1-13 (since there are 13) 🧵 alright, so first up is... the first one. opening b… Less
this year was @rlgrime's last halloween album release (13/XIII); time to bring it back all the way to the beginning. let's listen to and rate all of them from the start.
I'll rate each on a scale of 1-13 (since there are 13) 🧵
alright, so first up is... the first one. opening by rl stine the goat himself. you read that right - RL STINE, the book writer, someone I grew up with during 7yo-14yo reading his work.
anyway - now at the time of writing this, I'm listening through Halloween I, then I'll listen through Halloween II, and make a new tweet in this thread with the updated list...
And so on, and so on, until all 13 are listened to and rated.
RATING SCALE
(1st place is best, 13th place is worst):
1st: Halloween I
2nd: NA
3rd: NA
4th: NA
5th: NA
6th: NA
7th: NA
8th: NA
9th: NA
10th: NA
11th: NA
12th: NA
13th: NA
this seems insane. i do 15 grams of coffee per 400 grams of water. am i the retarded one or?...
wait a second, we're supposed to check the code? it just works for me every time, and when it doesn't, i tell the llm it doesn't work, and keep doing that until it works. @yacineMTB
I pray this to be the case, and not the alternative nightmare scenario: we're so in our internet culture bubbles, we think this to be true
I hope Brian is right here!
@brian_armstrong
This seems like the first election where new media has fully flipped traditional media. Long form podcasts, X/social, prediction markets etc deciding this election. Also holding traditional media accountable. Happened gradually then suddenly.
woah - it would be interesting to see what an algorithm based on "verified engagement only" would look like. i'd love a toggle for that
@doganuraldesign
X should show verified engagements.









