mirror of
https://github.com/ZHANGTIANYAO1/teamspeak-music-bot.git
synced 2026-10-01 20:42:50 +08:00
Extraction always used Matroska (.mka) because it takes essentially any audio codec. That is right for most codecs but wrong for AAC: MP4 records the AAC encoder priming (the ~1000 warm-up samples every AAC encoder emits) in an edit list, and the edit list does not survive into Matroska. The remuxed track then decodes ~23 ms longer than the source, with the priming samples played at the head instead of discarded. Measured on a 5s 640x480 fixture: source audio decodes to 962980 bytes of PCM, the .mka to 967440 — 4460 bytes / ~23 ms extra, peaking at -66 dBFS. Inaudible in practice, but it also puts the track fractionally out of step with its own reported duration, for no reason. Pick the container by codec instead: aac -> .m4a (keeps the edit list), everything else -> .mka as before. If the preferred container refuses the codec, retry into .mka before falling back to keeping the whole video. AAC is worth the special case because mp4 / mov / m4v — what people actually upload — almost always carry it. Adds the strongest available test of the "lossless" claim: decode the audio straight out of the source mp4, decode the stored extract, assert the PCM is byte-for-byte equal. Forcing .mka fails it with exactly the 4460-byte delta. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>