A trillion-parameter AI model is not actually as big as that number makes it sound

Started by Rosie_81, Yesterday at 01:36 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: A trillion-parameter AI model is not actually as big as that number makes it sound   Views(Read 58 times)
Active members in this topic:
Rosie_81(1)

Rosie_81

Modern AI model announcements love to lead with an enormous parameter count, a trillion parameters, several trillion parameters, numbers so large they stop meaning much intuitively. What a lot of coverage leaves out is that many of the biggest models today do not actually use their full size for every single task, thanks to an architecture called mixture of experts, which quietly changes what that headline number actually represents

The idea works something like a hospital. Instead of one single doctor trying to treat every possible patient walking through the door, a hospital has specialists, a heart doctor, a skin doctor, a brain doctor, and a receptionist whose entire job is routing each patient to the right one. A mixture of experts model works the same way, it is built from many smaller specialized sub-networks, and a small routing component decides which few of them actually need to activate for any given piece of text, rather than running the entire giant model every single time

This is why a model card can list something like 1.6 trillion total parameters right next to 49 billion active parameters without it being a typo, the total number describes the model's full stored knowledge, while the active number describes what actually gets used for any one specific answer. This is the trick that let AI models keep growing dramatically larger without becoming proportionally slower and more expensive to actually run, decoupling how much a model knows from how much computing power any single question actually costs

Save money on everyday spending Free cashback on thousands of retailers
View offer