3 Comments
User's avatar
Dylan Black's avatar

I love articles that give me good heuristics and mental models of complex fields. Great job!

Auspicious's avatar

Great summary of the state of scaling!

Enon's avatar

"We can make an AI that can do anything so long as we can specify it" Very big if true. As you noted early in your article, but set aside in the footnotes data varies widely in quality and meaning. Great gains can be made from filtering out the junk rather than adding more junk. Data also varies widely in type, the CRC handbook is not the Encyclopedia of [foo] is not sensor data is not Gene Wolfe is not Reddit comments. Lumping all data together is like lumping all matter together, only worse. Information is precisely that which is not fungible.

The real challenges are cross-correlating all the data after it's filtered, filtering it further, generating valid and useful new data /implied/ by existing data (not fake data resembling existing data), verifying it, and integrating it into a comprehensive world-model (not just physics but psychology and life in general).

World models are the essential missing ingredient in AI, without them it csn't ever be anything but BS, models of symbol strings rather than the world. Algorithmic scaling laws in many fundamental areas show that perfect simulations are always going to be infeasible, but other work shows that good enough simulations are usually possible for many purposes even for NP-complete problems. Still, there is always a lot it is easier to try experimentally than to simulate, and a lot of crucial things that just can't be predicted.

Other relevant things, not writing up at length here:

* I have a 15,000 volume electronic library hand-selected and verified by a top engineer for the purpose of restarting technological civilization and as a reference for creating self-reproducing manufacturing systems. This should be useful for AI training. Just how incredibly difficult and expensive it will be to properly tokenize with all the tables, graphs, equations and illustrations proves to me just how poor the utilization of data must be in training frontier models, all the equations in all the PDF textbooks and papers they used are turned into gibberish, no one did any verification or editing; slop in, slop out.

* Training AI to run such future automated, general-purpose manufacturing systems should be a prime goal, we aren't close yet. (The economic organization of such systems is even more important: ownership of the parts of the manufacturing system must be widely distributed and concentration prevented to to avoid demand collapse from former workers no longer getting paid and so no longer being able to buy the products being made; internal, automated markets are essential for such an "economy-in-a-box" to work. Widely distributed, non-concentrated ownership of AI token-processing work itself is needed to avoid automation causing demand destruction, hyperconcentration of assets and income, widespread destitution, and economic collapse.)

* The real data that we can easily get much more of is objective data, sensor data. My just-started Skywatch project will provide such data, using over ten thousand small telescopes per ~$30M observatory to watch the whole sky constantly at 30fps, with distance and velocity information out beyond geosynchronous orbit, with many observatories around the globe together able to triangulate out to over 500M km, each observatory producing 20TBps raw data, distilled into a constantly updated model plus residual surprising information. Needs Starlink for remote sites, good use of spare bandwith, aids AI and space development, essential for planetary defense, lots of high-quality open science data, $600M-$10B tax-deductible - Elon, are you listening?

** Seeking funding for any/all 3 projects.