Browse multilingual, audit-ready audio corpora with quality metadata, provenance, and training-ready formats.
phonarch / pipeline-feed
Approved supplier logos will appear here once provided.
Each stage runs as its own agent on a shared schedule — scout emits, rights licenses, rubric scores, shard builds. A slug that fails a bar stays in the queue; a slug that clears every bar lands in a release.
rolling 24h
Indexes candidate recordings across open broadcasters, podcast feeds, and consented contributor pools.
p50 6m
Reads consent docs, verifies chain of title, attaches jurisdiction-specific license metadata on every file.
p50 41s / clip
Scores against task-specific rubrics — WER threshold, SNR floor, accent coverage, and demographic balance.
p50 12s / shard
Writes training-ready shards aligned with Whisper, SeamlessM4T, Parakeet, and USM ingestion formats.
Scoring is rubric-driven and task-specific. The thresholds are configurable per release, but every score is reproducible from the shard manifest — cited, not asserted.
| Rubric | Floor | Method |
|---|---|---|
| Word error rate (WER) | ≤ 8% on paired transcripts | Cross-checked against three independent ASR passes; outliers flagged for human review. |
| Signal-to-noise (SNR) | ≥ 18 dB on 95th-percentile frames | Per-channel noise floor measured across full clip, not segment heads. |
| Accent coverage | ≥ 14 regional variants / target language | Demographic tile maps ensure no single accent dominates the cohort. |
| Demographic balance | Reported per shard, audited per release | Speaker age, gender, and dialect are sampled against the population your model targets. |
| Provenance chain | One signed manifest / clip | Every audio file carries a tamper-evident trail from consent capture to packaged shard. |
| Conversation realism | Topical + turn-taking plausibility | Call-center and conversational subsets filtered for natural disfluency, interruption, and overlap. |
Drop-in ready for the benchmark speech stack. Each release is shaped for ingestion as-shipped — no studio-side wrangling, no bespoke ETL, no surprise schema rotations between releases.
Segmented JSONL + manifest
Speech-text paired shards
Low-latency chunking
Long-context pretraining
We work with teams whose data flows are reviewed by counsel and whose datasets must survive a recurring vendor audit. If that's you — read on.
Scaling open-source speech models needs multilingual corpora with auditable rights and quality bars that hold up across releases.
Localizing real-time voice agents requires conversational corpora in the dialects your callers actually use — not studio-clean tidied samples.
Expanding language coverage on ASR products means WER-controlled, accent-balanced datasets with paired ground truth you can defend in a procurement review.
Sourcing studio-quality recordings alongside conversational corpora, priced at market-standard rates that survive a procurement audit.
Continuous refresh cycles keep model iteration velocity as the audio vertical outpaces text and image inside the broader AI training dataset market. The paper trail is part of the product, not a paid add-on.
Request the audit packetGDPR
Right-to-erasure runs through the same pipeline that scored the clip. Erasure is reproducible from the audit trail.
HIPAA
Healthcare call-center subsets ship stripped of identifiers before they touch a scoring agent, with a manifest entry that pins the strip step.
Speaker attribution
Consent capture, jurisdiction tag, license tier, and clip-level rubric scores are signed together — one tamper-evident manifest per shard.
Continuous refresh
Weekly drops replace flagged or stale clips without re-running the full pipeline — relicense, rescore, repackage in place.
Replies within one business day. We'll sign an NDA before sharing sample shards.