01Supply
Internet-scale
web data for
AI Labs.Business Intelligence.Multimodal Research.
We collect any video, audio, text, image, product and market data on the public web and deliver it to you as a finished consignment. No integration, no infrastructure to run.
| Ref | Corpus | Modality | Volume | Region | Status |
|---|---|---|---|---|---|
| CN-48231 | reviews-global | text | 2.4 TB | eu-central | packing |
| CN-48229 | video-index | video | 8.1 TB | us-east | in transit |
| CN-48226 | commerce-eu | text | 640 GB | eu-west | delivered |
| CN-48224 | audio-corpus | audio | 3.7 TB | us-west | delivered |
| CN-48221 | listings-apac | image | 1.2 TB | apac-south | delivered |
| CN-48218 | news-archive | text | 5.6 TB | eu-central | delivered |
| CN-48215 | social-stream | video | 9.3 TB | latam | delivered |
250Gbps
Sustained throughput
106
Markets covered
85
Languages supplied
10M+
Residential endpoints
02What we supply
Two supply lines
Both arrive cleaned, de-duplicated and mapped to your schema. Pick the one that matches what you are training on, or take both.
Multimodal Data
Video, text, image and audio pulled from the public web through a single supply line, filtered and compliance-ready. Advanced filtering and synthetic annotation are available on request.
- Modalities
- Video · Text · Image · Audio
- Formats
- 12
- Languages
- 100+
- Max file size
- 50 GB
- Annotation
- On request
- Delivery
- S3 · GCS · Direct
Datasets
Curated corpora scraped from anywhere on the public web, de-duplicated and structured to your schema. Delivered once or on a recurring schedule.
- Coverage
- 106 markets
- Formats
- JSON · CSV · XML · Parquet
- Refresh
- One-time or recurring
- De-duplication
- Included
- Provenance
- Per record
- Delivery
- S3 · GCS · Direct
03Coverage
Every modality on the public web
Provenance on every record
Source URL, capture time and collection method travel with the data, so your legal team can audit any row.
De-duplicated before handover
Near-duplicate detection runs across the whole consignment, not per batch, so you are not paying to train on the same page twice.
Annotation to your schema
Synthetic labels and segmentation are applied to your spec before delivery rather than left for your team to do.
04How supply works
Three steps, no engineering
You never touch a crawler, a proxy pool or a queue. You describe what you need and we hand over the finished data.
- 01
Scope
Tell us the sources, modalities, volume and schema. We come back with what is collectable, what it costs and how long it takes.
Typically 2 working days - 02
Sample
We collect a representative slice and hand it over in your target format so your team can run it through evaluation before committing.
Free of charge - 03
Delivery
The full consignment lands in your bucket, once or on a recurring schedule, with provenance attached to every record.
S3 · GCS · Direct
05Next step
See the data before you buy it
Send us the brief. We will come back with a scoped quote and a free sample consignment in your target format.