the dialogue you mixed is not the dialogue they hear
a chain of reasonable decisions ends at the sofa with the voice buried. the film is about the chain. this issue is about your end of it: how to mix dialogue so it survives a stereo fold-down, a flat panel and an ordinary pair of ears, with the numbers i use, including one my own films failed on.
a director sends me a rough cut and asks why the voice feels small. it is not small. on her monitors it is right there, sitting exactly where she put it. then she watches it on the television in her kitchen and half the lines go under.
the film this week is about why that happens. three loudness targets that disagree, a metadata note nobody checks, a fold-down coefficient in a table, a speaker that is a few centimetres deep. every link is published and measurable. the one thing everybody says, that it has got worse, is the one thing i could not find a measurement of.
this issue is the other half. i have mixed dialogue for videos and for a couple of games, and most of what i know about it i learned by getting it wrong in ways that were easy to measure afterwards. so here is what i do now.
the frame
you do not mix the dialogue. you mix the ratio.
the voice on its own is fine. everything that arrives on top of it between your room and theirs is the problem, and each of those things is a ratio you can measure before you send the file.
the ratio, and the number it failed on
for months my spec for the music bed under narration said “duck the bed 6 to 9 dB”. that is a process. it says what the fader did. it says nothing about what the listener gets.
then two people wrote in about the ai video i did last week on the channel to say the music was too loud. i went back and measured the delivered speech to bed ratio on every film with a stem. four out of five sat between 7 and 11 dB. the one people complained about had been ducked 6.3 dB, inside the spec, and sat at 7.2 dB. the one film nobody complained about sat at 28. and sorry for that loud music to you who wrote the comments.
so the spec changed from a process to an outcome. the bed now has to sit at least 15 dB under the speech, broadband, on the delivered file. and at least 20 dB under it between 1 and 4 kHz, because that is where the consonants live, they are quiet compared to the vowels, and it is also where a flat panel is weakest. a gate measures both numbers on the mix and the isolated bed stem before anything ships.
and this week that gate passed a mix it should have failed: it was reading the bed stem from before the output gain, so the ratio it printed was 2.4 dB better than the one in the file. a reviewer reading the code found it. a checker is only as honest as the thing it is pointed at.
the fold-down, on your desk
if you mix in 5.1, the standard puts your centre channel into the left and into the right at 0.7071, which rounds to minus 3 dB, and then adds the surrounds into the same two outputs at the same coefficient. the voice goes down as it is folded in, and everything that used to have its own speaker lands on it.
so fold it yourself before delivery. render the stereo downmix with those coefficients, then play the dialogue alone and the downmix back to back. if a line goes under in the fold, something in the surrounds is sitting in the voice’s band. fix it in the mix. the downmix is arithmetic and will do the same thing again.
if you mix in stereo, do the equivalent: sum to mono and listen for the lines that thin out. and check the phone. the phone is where the most people will hear it.
mix it louder than your ear wants
this one is uncomfortable and it is well measured. when researchers tested trained engineers against ordinary listeners on the same material, the ordinary listeners wanted the dialogue about four loudness units louder than the engineers chose. when the background was another voice, the gap got much wider: the engineers set the voice 11.5 to 17 LU above it, and listeners over fifty seven wanted 20 to 30.
that is not a mistake by either group. it is two populations with different ears. your ears are the trained ones, and the default mix has to serve the other group. so when the dialogue sits exactly where you like it, push it up a little and leave it there. it will feel slightly wrong on your monitors. it is right on their television.
the note that travels with the file
if you deliver for broadcast or streaming, two things are yours and nobody downstream will check them. the loudness target belongs to the deliverer, and there are three current ones that disagree, so ask which one before you normalise. and the dialnorm value is a note attached to the audio that tells the decoder how far to turn everything down. if the note is wrong, the whole programme lands at the wrong level and nothing in the chain notices. write it from a measurement, not from the last project’s template.
the part that is also about work
every link in that chain was set by a reasonable person who never met the next one. the loudness body, the encoder, the standards committee, the panel designer, the mixer. each did their job. the voice still went under.
that is what most handoffs look like. the failure lives between the steps, in the thing nobody owns. so the useful question is “what would i have to measure at my end so the next person could not get it wrong”. a ratio on the delivered file is that measurement. it is boring, it takes a minute, and it is the only part of the chain you control.
what this changes
take a mix you have already delivered. measure the speech to bed ratio on the delivered file, broadband and in the 1 to 4 kHz band, over the spoken lines only. if either number is under 15 and 20, that mix is quieter on a television than it was in your room, and now you know by how much.
then fold it down, sum it to mono, play it on your phone at a conversation level. listen for the line that goes.
from the studio
oh, and a short one on OPEN. still working hard on it. it went through more blind listening sessions in one week than any plugin i have made, and something should land soon for everyone, not only the beta group.
what does your television do to a mix you know well? reply, i collect these.
jonas
more field notes
Sep 15, 2026
·sound science
your loudest master never arrives loud
streaming turns your master down to a target before anyone hears it, so mastering louder buys nothing but lost dynamics. and the loudness war mostly did not do what people think: measured across five decades the spread between loud and quiet did not narrow. what came down was the transient crack, by about a decibel and a half. so the thing to watch is heavy limiting and clipping, more than compression as an idea, and the move is to master until it is finished and let the platform set the number.
Sep 1, 2026
·sound science
a row of spikes nobody trained in
generated music leaves a mark, and it has nothing to do with taste. it is a row of evenly spaced spikes, put there by an operation that is also sitting inside half the plugins you own. which gives you a rule you can use on any analyzer: evenly spaced peaks came from a process, musically spaced peaks came from the music. plus a second tell the paper does not cover, and a note on why neither one is worth trying to EQ.
Aug 25, 2026
·sound science
nobody publishes where a phone speaker stops
two films, a few days apart, and the same thing happened twice. i went looking for the frequency a phone speaker gives up at so i could mix against it, and it isn't published anywhere. then i sat down to check the tuning arithmetic and found the vallotti preset in almost every synth is named after the wrong person. go and check a number, and it's either not published or it's filed under someone else's name.