Skip to content
Corpshore Polska

2026-09-08

AI data annotation in the EU: why labelling training data is a GDPR question too

Training and fine-tuning AI models requires annotated data, and that data is often personal. Choosing an annotation vendor outside the EEA creates the same transfer problem as any other outsourcing.

Data annotation sounds purely technical: someone labels objects in an image, rates a model's response, or transcribes a recording. In practice most training sets contain personal data, faces in images, voices in recordings, names and health details in customer service transcripts, message content in language model fine-tuning data. When personal data reaches an annotator outside the European Economic Area, the same transfer obligation arises as with any other processing arrangement: a transfer mechanism, an assessment of the destination country's law, supplementary measures. This is general information, not legal advice.

The argument is identical to any other processing outsourced inside the Union: when the annotator operates inside the EEA, the transfer mechanism, the transfer impact assessment and ongoing adequacy monitoring stop applying, because there is no cross-border transfer to legalise. For training data work this matters more than usual, because annotation sets tend to be large, reprocessed across successive fine-tuning rounds, and retained for a long time for model reproducibility, which extends the exposure window rather than shortening it.

Good annotation inside the EEA looks different from cheap offshore annotation, independent of legal jurisdiction. Inter-annotator agreement tracking, native-language quality, especially in Polish, Ukrainian and German, an audit trail showing who labelled which record and when, and documentation ready for a data protection impact assessment. Corpshore Polska runs this work as part of its data annotation service, within a group ranked fifth of fifty AI outsourcing companies worldwide by Outsource Accelerator.

Choosing an annotation vendor deserves the same questions as any other processing arrangement: where is the data physically hosted, where do the people doing the annotation sit, does the vendor use sub-processors outside the EEA, and what does deletion or return look like at the end of the project. A provider with an EU registration that routes annotation work through a platform or crowdworkers outside the EEA has introduced a transfer through the back door, exactly as in any other outsourcing arrangement.

Frequently asked questions

Does AI training data annotation fall under GDPR?
Yes, if the labelled data includes information about an identified or identifiable person, which covers most sets involving images, voice recordings, call transcripts or text messages. Anonymisation removes the obligation only when it is irreversible, which in practice happens less often than buyers assume.
How do you choose a GDPR-compliant data annotation vendor?
Check where the data is hosted, where the annotators sit, the full sub-processor list, and the deletion process after the project. A provider inside the EEA with no non-EU sub-processors removes the transfer problem entirely.
Does annotation in Poland cost more than offshore?
Usually yes, on a per-seat basis. The difference is justified where personal data cannot leave the EEA, where native-language quality in the target language is critical, or where compliance documentation needs to be audit-ready rather than assembled after the fact.

Let us talk about your team in Poland

We respond to every enquiry within six hours. Book a call or request a proposal.

Looking for a role at Corpshore? See our open positions. Open roles