Blog by Tajh L. Taylor: “The public discourse about new data center construction has reached a fever pitch. Local and state governments are enacting moratoria on new data center construction, because the political reality is that citizens have decided they don’t want to endure escalating energy prices and water consumption for the benefit of a few massive corporations datacenter bans. AI companies themselves are going into so much debt to build out computing infrastructure that it is possibly affecting 30 year Treasury yields, which are at their highest point since 2001. As is always the case, when public discourse gets this polarized, the sensible path forward becomes harder to see.
I feel the need to keep repeating this: not all AI needs massive amounts of water or power, nor does it even necessarily need to be hosted in massive data centers powered by highly polluting generators futurism. Lots of good and useful AI applications are being deployed that run on more ordinary-scaled hardware, or even small devices “at the edge” – meaning in your office, on the factory floor, in your home, on your phone. The application areas are varied and include medical diagnosis, voice assistive technology, sorting produce, fraud detection, et cetera. Many of these application areas predate the current generative AI boom, but have benefited from the substantial investments into AI research and product development.
Moreover, not all AI labs working on gen AI are working on hyperscaled intelligence. Some of the labs working on the frontier of small and efficent AI:
- Rootcomputer builds and studies models with 7B or fewer parameters. That’s small enough to run on many consumer devices such as phones and laptops.
- Sakana.ai is doing research in a few different directions, including on efficiency of key-value memory in training while keeping compute efficiency low.
- Liquid.ai works on device-native foundation models aimed at deployment on phones, laptops, cars, and other embedded hardware.
Some of the bigger and more well-known companies also work on this:
- Qwen is Alibaba’s AI lab, and they have produced a series of increasingly powerful small models that have taken almost all the recent attention in the AI self hosting world.
- Apple ML Research works on intelligence specifically constrained to (Apple) phone/laptop hardware
It’s not a coincidence that you don’t hear or see much about this work. The loudest voices in this conversation aren’t interested in these paths, and they aren’t doing themselves any favors in persuading anyone to take their positions who isn’t already aligned. Like so much discourse these days, it is more of a sorting exercise than a discussion to develop thoughts and change minds.
Here’s what will break the logjam: as the frontier of the most powerful AI continues to advance, fewer use cases need frontier capability. And as companies deploying AI have learned in the past few months, it isn’t a smart use of budgets to max out token spend on the most expensive models for every task. As long as frontier model advancement continues to advance inference costs (which seems to be the case as parameter sizes grow at the largest scales), there will be incentive to carefully decide which tasks can be done by which models. At some point, the irrational exuberance around AI investment and spending will give way to rationality – but to distort the Keynes saying, it can remain irrational much longer than we can bear the cost…(More)”.