Annodex logo AnnodexNotes for people who make media
Games

How Online Game Voice Chat Audio Is Cleaned Up in Real Time

Online game voice chat should be a disaster. Players use cheap headsets, laptop microphones and phones. They sit next to loud fans, mechanical keyboards, televisions and barking dogs. Their internet connections drop packets.

Abstract illustration for How Online Game Voice Chat Audio Is Cleaned Up in Real Time

Online game voice chat should be a disaster. Players use cheap headsets, laptop microphones and phones. They sit next to loud fans, mechanical keyboards, televisions and barking dogs. Their internet connections drop packets. And yet most of the time, you can understand your teammates. That is because a surprising amount of audio processing happens between their mouth and your ears, in real time, in a few milliseconds.

I produce podcasts, where we clean audio carefully after recording. Game voice chat has to do similar work instantly. Here is how it is done.

Capture: the microphone and the gate

The first step happens on the player's device. Most voice chat systems offer two ways to decide when to transmit: push-to-talk, where a button opens the microphone, and voice activation, where the software listens and transmits when it detects speech.

Voice activation uses a noise gate, which stays closed until the signal crosses a threshold. A badly set gate either cuts off the start of words or lets keyboard clicks through. Modern systems use voice activity detection based on machine learning, which recognises speech patterns rather than just volume, and handles noisy environments much better.

Noise suppression and echo cancellation

Next comes noise suppression. Traditional methods estimate the background noise spectrum and subtract it. Newer neural network approaches, such as the RNNoise library used in several apps or NVIDIA's RTX Voice and Broadcast tools, learn the difference between voice and noise and can remove fans, typing and even some background voices with impressive results.

Echo cancellation solves a separate problem. If a player listens through speakers, their microphone picks up the game and other players' voices, sending everyone an echo of themselves. Acoustic echo cancellation subtracts the known output signal from the microphone input. It works well, but headphones still work better, the same lesson podcasters learn, as covered in the cheapest upgrades that actually improve podcast audio.

Level control: automatic gain

Some players whisper, some shout, some sit far from their microphone. Automatic gain control adjusts levels so everyone sounds roughly equally loud. Without it, one excited teammate can deafen the rest of the squad. With too much of it, quiet background sounds get pumped up during pauses. Good systems strike a balance and usually let players adjust their own input and each teammate's volume separately.

Compression for the network

Voice has to travel over the internet with as little delay as possible. Most modern game voice systems use the Opus codec, which was designed for exactly this purpose. Opus can run at low bitrates, often around 16 to 40 kilobits per second for voice, while still sounding clear, and it adapts to network conditions on the fly.

Voice packets usually travel over UDP rather than TCP, because a late packet is useless for a live conversation. When packets are lost, the codec uses packet loss concealment to fill tiny gaps with plausible audio, so you hear a smooth voice instead of clicks and dropouts.

Jitter buffers and mixing

On the receiving side, packets arrive at uneven intervals. A jitter buffer holds them briefly and plays them out at a steady rate. A larger buffer handles worse connections but adds delay. Systems adjust the buffer size dynamically to keep latency as low as possible.

Finally, voices are mixed with game audio. Many games apply ducking, lowering game sounds slightly when someone speaks, and some place voices in 3D space so proximity chat sounds as if it comes from the speaker's character. On the platform side, providers run large voice server networks that route and mix these streams; when I compared voice quality across several online game platforms, the clearest results came from services with well-tuned regional voice servers, such as those used by an online gaming network around ankertoto where latency stayed low even across busy evenings.

Is AI noise removal always a good idea?

It is tempting to switch every noise suppression feature to maximum. I think that is usually a mistake. Aggressive processing can make voices sound robotic, cut off quiet syllables and introduce strange artefacts when two people talk at once. Stacking several suppressors, such as one in a headset app, one in the operating system and one in the game, often makes things worse.

Use one good suppressor at a moderate setting, a decent headset microphone positioned close to the mouth, and push-to-talk in very noisy rooms. Simple habits beat extreme processing.

Safety and moderation

Voice chat processing increasingly includes moderation. Some games now use automated systems to detect harassment in voice channels and flag it for review. Combined with mute and report tools, this aims to make voice chat a place more players feel comfortable joining.

For players who cannot use voice at all, speech-to-text and ping systems are vital, as described in how captions help online game players follow along. More in our Games section.

TB
Tariq Bellweather

Tariq learned audio at a community radio station and now produces interview podcasts. He writes about microphones, rooms and editing speech so that it sounds natural, and he has strong opinions about loudness.

More posts by Tariq

More in Games