Right. When we design institutions like these, what matters most is to think of the people who have no time to take part, and the people whose language was never in the mainstream data, so that even when they tell a language model they have been harmed, the model cannot understand them. What they face is what is called epistemic injustice : the more they speak, the more others take them to be talking nonsense, or take them for people not worth hearing. As automated judgment spreads, that gulf only widens. So suppose you say at the outset that this is what your system is for. Mozilla, for instance, has a project called Common Voice , where any community, Indigenous communities among them, can donate for itself: you read aloud the everyday words your people wrote for one another, and you donate those voices. They enter the Mozilla Data Collective , something like a data-production cooperative. And the so-called revenue share is that the small models it trains, the post-trained ones, take care of you first.