Every model team hits the same wall eventually. The English data is fine, the pipeline works, and then somebody asks for Swahili, or for a Gulf Arabic dialect, or for eight hours of labelled operating-theatre video, and the vendors who were easy to buy from go quiet.
How do you get training data in a low-resource language?
You recruit speakers, you do not scrape. For a language with little text on the web there is no corpus to buy, so the data has to be produced: native speakers writing, speaking, transcribing and ranking, screened for the dialect you actually need rather than for the language name on their profile.
The screening is the whole job. "Speaks Arabic" covers varieties that differ more than Spanish and Portuguese do, and a model trained on the wrong one sounds fluent and wrong to every user you were building for. The same is true of Swahili in Nairobi against Swahili on the coast, and of every language with a standard written form and a spoken one people actually use.
- Screen on the variety, not the language Region, first language, where they were schooled. A single question about a word that differs between varieties sorts most of it in one step.
- Pay properly and locally Rates that work in one economy do not in another, and underpaying is how you end up with fast, careless annotation and a quality problem you cannot see.
- Measure agreement from day one Inter-annotator agreement on an overlapping sample is the only early signal that your guidelines are ambiguous rather than your annotators being wrong.
- Write the guideline against real examples Guidelines written in the abstract produce disagreement at the first genuinely hard case, which arrives on day one.
- Keep a held-out set nobody in the pool has seen It is how you tell a pool that has learned the task from one that has learned the annotator.
Can you label live and recorded video?
Yes. Action labelling, object tracking across frames, scene and safety classification, transcription with timed captions, and event marking on live streams. Video is slower per hour than text by a wide margin, which is the thing to plan a budget around rather than be surprised by.
The rate depends almost entirely on density. A fixed camera on a quiet room might run close to real time. Multi-object tracking through a busy scene can take fifteen to thirty minutes of annotator time per minute of footage, and no honest quote can be given for video without somebody watching a representative sample of yours first.
Illustrative of the ratios rather than a quote. The point is the scale of the difference: an hour of footage needing dense tracking is most of a working week, and a budget built on a text-labelling rate will be out by a factor of twenty.
Growth is a process. Nothing here happens overnight, and anybody promising you overnight results is lying to you.
Price a pilot from $100What is RLHF data, in practice?
People ranking model outputs against each other and saying which is better, with a written reason. The ranking is the data. The reason is what lets you find out later whether your raters understood the task or were all quietly optimising for length.
It is more sensitive to rater quality than any labelling task, because there is no ground truth to check against. A preference dataset built by a pool nobody calibrated is a dataset that teaches a model the pool's habits, and those habits are invisible until the model has them.
How do you know the data is any good?
Three things, reported weekly rather than at the end: inter-annotator agreement on an overlapping sample, throughput per annotator, and the rate at which items are sent back. A vendor who reports only volume is reporting the one number that cannot go wrong.
Ask for those three before you sign anything. If they cannot be produced, the pool is not being measured, and a pool that is not being measured is one whose quality you will discover during training rather than during collection.
How small can a first project be?
Small enough to be a test rather than a commitment. Our entry pilot is $100: a defined slice of your task, run by a screened pool, returned with the agreement score, so you can look at real labelled data of your own before deciding anything about volume.
That is the right first purchase for the same reason it is elsewhere on this site. Guidelines that read clearly in a document turn out to be ambiguous the moment two people apply them to the same hard example, and finding that out on a hundred items costs a fraction of finding it out on a hundred thousand.