Two very different reasons to go mono
Almost everyone who converts stereo to mono is solving one of two problems, and they want opposite things from the tool.
The first is a size or compatibility problem. A podcast episode, a lecture recording, an audiobook chapter or a voice memo has no meaningful stereo content, so carrying two channels is wasted space. Folding them together is tidy housekeeping and the result should sound the same.
The second is a broken recording. An interview came back playing in one ear only, because a lapel microphone was plugged into a single input and the other channel recorded silence. Here the goal is rescue, and the naive fix makes it worse. That is why this page has three modes rather than a single button.
Converting a file
- Drop an MP3, WAV, M4A, OGG or FLAC onto the box. One file at a time.
- Choose a Downmix mode. Average is the default and is correct for real stereo.
- If the audio plays in one ear only, pick the mode matching the ear where you can hear it.
- Run it and download the result, which is an MP3 named after your file with a mono suffix.
Why averaging is the wrong repair for a dead channel
The instinct with a one sided recording is to combine both channels, since combining is what going mono means. Work through the arithmetic and the problem is obvious.
An average takes half of each channel. If the left channel holds your interview and the right holds nothing, the result is half of the interview plus half of nothing, which is the interview at roughly half its former level in a single channel. The audio is now in both ears, which feels like progress, and it is six decibels quieter than it was, which is not.
Taking one channel and discarding the other keeps the recorded level intact. The trade is that anything genuinely present on the discarded side is gone, which is exactly what you want when that side was silence and exactly what you do not want for music.
Why the output is always an MP3
Some conversions on this site copy the compressed data straight through without touching it. This one cannot. Combining channels means arithmetic on the actual audio samples, which means decoding the file, mixing, and encoding again.
Given a re-encode is unavoidable, the choice is what to encode to, and the tool standardises on MP3 at 192 kbps rather than trying to match your input. That rate is comfortably transparent for speech and holds up well for music, and MP3 plays everywhere without argument. The consequence to be aware of is that a lossless input does not stay lossless. If that matters, keep the original alongside the mono copy, and if you specifically want a controlled MP3 conversion from a WAV, WAV to MP3 is the tool built for that with a bitrate you choose.
What to do after, or instead
If the underlying complaint is that a recording is too quiet rather than one sided, converting to mono will not help and Normalize Audio or Boost MP3 Volume will. If the problem is your hardware rather than the file, confirm it first with Speaker Test, which drives each channel independently and settles the question in about ten seconds.
Mono is also the sensible input format for anything that reads speech rather than plays it, so converting first is a reasonable step before running a file through Transcribe Audio to Text. To separate a song into its parts rather than merge its channels, Vocal Remover is a different operation entirely. The rest are on the audio hub.

