AI UNIT 2 • STAGE 3 OF 5
From a handful of examples to billions, and whose voices made the cut
In Stages 1 and 2 you taught the AI with a few examples. A real AI like the one in this sandbox learned from a pile so big it's hard to picture: billions of pieces of text, most of it collected from the open internet, websites, articles, forums, books.
That's its training data. Everything it can say, it can say because something like it appeared, over and over, somewhere in that pile.
Put it to the AI directly, and pay attention to whether it admits the gaps.
Now connect Stage 2 to this. If the pile is mostly writing from some communities and very little from others, the machine learns those imbalances. And for Native peoples, much of what's in the pile was written about your communities by outsiders, not by your communities.
An AI can't be richer than its pile. Where your nation's own voices are thin online, the machine fills the gap with outsiders' words, old sources, and stereotypes, exactly the "garbage in" problem from Stage 2, at a massive scale.
In your Field Notes, describe what's in the pile in your own words, whose voices are well represented and whose aren't, and name one thing about your community that's probably thin online.
You've gone from teaching a machine three examples to understanding a pile of billions. Stage 4 asks the question that's been building all along: who decides what goes in that pile?
Map what's in, and what's missing:
Saved automatically.