Skip to content

Commit b40497c

Browse files
author
Michal Warda
committed
Record YouTube dataset validation
1 parent 26d5015 commit b40497c

1 file changed

Lines changed: 8 additions & 0 deletions

File tree

docs/pipeline/PROGRESS.md

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -114,6 +114,14 @@ was an example batch size, not a workflow limit.
114114
from an already-failed job whose source objects never reached S3. The current
115115
AudioSet set reports 98 stereo and two mono originals with exact codec,
116116
sample-rate, bit-depth, and bitrate facts.
117+
- Local general-YouTube set `youtube-random-1000-20260714` contains exactly
118+
1,000 ten-second WAVs from 1,000 unique direct YouTube videos selected across
119+
288 randomized mix-biased queries. Its independent audit passes with no
120+
failures: every output is stereo PCM16/48 kHz, all source streams are stereo
121+
at 44.1 kHz or better and at least 120.084 kbps, and every loudness, silence,
122+
true-stereo, clipping, duration, uniqueness, existence, and SHA-256 gate
123+
passes. The audio payload is 1,920,078,000 bytes; no AudioSet record metadata
124+
or AudioSet candidate source was used.
117125
- Live mono-filter job `5ec26b1bb28b420a9245e21d475d26be` completed with
118126
`non_stereo_input`, one input channel, zero chunks, zero stems, and zero model
119127
tasks. A live `targets=voice` API canary produced only voice and residual

0 commit comments

Comments
 (0)