First impression: Third-party benchmarks or gtfo. Personally, I've never heard of either of these companies before. We're just supposed to take their word that they've matched the best models on the market?
Sakana describes their model as a "Orchestration Model." Does that mean that it's actually a bunch of different models glued together?
Is it actually that hard to make good models or is it just about the amount of resources you have to do training? (This is an actual question, I really don't know.) I'm sure it's not trivial but does it really take world class secret knowledge to build off of the known existing techniques? I feel like there's tons of low hanging fruit still to explore, and time and resources are the limiting factor.
Like Zuckerberg, top talent may not work with a polarising character if they disagree with his behaviour. Space focused talent don't have many choices aside from SpaceX but ai companies are a plenty and a top AI person can pick and choose.
I suspect that Grok has been ironically lobotomized by pressures to correct its political views.
Similarly, I could imagine the Gemini folks working in a significantly more complex corporate climate, with different parts of Google pushing for different capability focuses. They are only lagging behind less than a year, so it isn't too large of a gap yet.
That said, the fact that Anthropic is currently the top dog suggests that talent and execution is incredibly important. A year ago none of my normie friends new them, and when i suggested using Claude looked at me like when I recommend Linux.
That shouldn’t affect Grok’ coding ability. How often are people discussing politics with Claude code? Writing decent code is just hard and it’s not just Grok.
If training a good model requires talent then that’s the answer to the question this thread is trying to answer: is training a good model actually that hard?
Talent to do.. what? This could mean a lot of things.
Navigating astronomically huge fundamentally not so hard but still really tangly and hairy projects requiring both excellent short- and long-term vision in an overheated domain with angry people and lots of money is a skill all of its own.
You’d be quite surprised, I think. Fine tuning a model on one axis can have drastic impacts on another that as a human we would expect to be completely unrelated.
I have never seen anyone argue that this cannot be overcome with more high quality RLVR data.
The practical reality is that the Chinese and American models might have very different politics. But the most relevant factor in model performance is the quality and volume of training data, not ideology of the base model. Unless you are suggesting something very particular about the way Grok was neutered.
If they have a top team and the money then appears to be a matter of a year or two? And one startup mentioned is Japanese not Chinese so they won't be banned from buying US tech.
My impression is that the answer is yes, that it purports to dispense the glue on-the-fly in some kind of dynamic way rather than being some kind of new model-amalgam.
No, stop right there. Anything published by Anthropic implicitly is not third party. For it to be third party, the third party has to be the one publishing it.
When you're announcing a new model, typically, nobody else has benchmarked it yet, because it hasn't been released yet. You can still run 3p benchmarks on it and publish those results. If other parties later run the same benchmarks independently, and find major discrepancies, that would be a scandal.
And even if they didn't, they have a track record. Even if we did have benchmarks in this case I would still wait until people got there hands on it and formed a more holistic opinion.
Sakana describes their model as a "Orchestration Model." Does that mean that it's actually a bunch of different models glued together?