Benchmarking
51%75 posts analysés sur les 12 dernières semaines
Sur les 12 dernières semaines
Moyenne tous posts confondus
Ce mois-ci vs le précédent
75 posts analysés sur les 12 dernières semaines
Sur les 12 dernières semaines
Moyenne tous posts confondus
Ce mois-ci vs le précédent
Pente de progression (6 semaines)
Meilleure semaine : 13 juil. (242 likes moy.)
Andre Zayarni
Co-founder & CEO @ Qdrant. Search AI for humans, agents, and robots.
Ok, Elastic. You wanted it that way. You created a benchmark to show that your paid premium feature is superior to the open-source one, right? You published a benchmark compared to Qdrant, right? Oh, Honey, I Shrunk the …
Clem Delangue 🤗
Co-founder & CEO at Hugging Face
Big unlock for open-source AI inference: Hugging Face Transformers models can now run in vLLM at native speed, often matching or beating hand-written implementations. Until now, every new architecture often needed to be…
Steve Suarez®
Chief Executive Officer | Entrepreneur | Board Member | Senior Advisor McKinsey | Harvard & MIT Alumnus | Ex-HSBC | Ex-Bain
White House ran a quantum summit last week. IBM, PsiQuantum, D-Wave, Quantinuum all showed up. Commerce is putting $2B into 9 quantum companies. Everyone's watching that part. Real story's elsewhere. DARPA is runni…
Daniel Han
Co-founder @ Unsloth AI
My 2hr advanced AI workshop is out now! 🔥 I cover which LLM benchmarks to trust, reward hacking, kernels, closed vs open models & RL. Topics: 1. Open model progress 2. Which benchmarks should we trust? Benchmaxxing & ch…
DJ Kim
Lean Coach | Looking forward to the next chapter - eager for meaningful work in any form I Author of When Nike Met Toyota
𝗧𝗵𝗲 𝗠𝗼𝘀𝘁 𝗗𝗮𝗻𝗴𝗲𝗿𝗼𝘂𝘀 𝗘𝗻𝗲𝗺𝘆 𝗮𝘁 𝗧𝗼𝘆𝗼𝘁𝗮 𝗜𝘀 𝗧𝗼𝘆𝗼𝘁𝗮 When people hear the phrase "𝘿𝒆𝙛𝒆𝙖𝒕 𝑻𝙤𝒚𝙤𝒕𝙖," they usually assume it was aimed at competitors. It wasn't. In 2001, Toyota Cha…
Jan Beger
Our conversations must move beyond algorithms.
The same LLM that scores 92 on the US medical licensing exam scores 44 when you point it at real clinical notes. 1️⃣ A new benchmark called BRIDGE tested 95 language models on 87 real clinical tasks, in 9 languages, fro…
Ethan Mollick
Associate Professor at The Wharton School. Author of Co-Existence, coming October 20!
For the first time, the personalities and approaches of the leading models are diverging in significant ways, magnified by the fact that over longer task horizons these differences in judgement & approach are magnified. …
Robert F. Smith
Founder, Chairman and CEO at Vista Equity Partners
Right before the holidays last year, I had the opportunity to see back-to-back AI agent demos from more than 20 of our CEOs. They were ten minutes each, though I could have spent an hour with every team. Seven months ago…
Adam Holmgren
Co-founder & CEO at Fibbler | Building the future of Paid Ads Attribution for SMBs | 10+ years in B2B marketing
You can see every LinkedIn Ads number you have. You still have no idea if any of them are good compared to companies like yours. 𝗔𝗻𝗱 𝘆𝗼𝘂 𝗺𝗶𝗴𝗵𝘁 𝗯𝗲 𝘄𝗮𝘀𝘁𝗶𝗻𝗴 𝗺𝗼𝗻𝗲𝘆.. Here is the thing. There is no…
Ethan Mollick
Associate Professor at The Wharton School. Author of Co-Existence, coming October 20!
As a joke I prompted Codex "Build and run BenchBench, a benchmark of now good AI is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a good arXiv paper." …