Synthetic speech detection performance considerations
How does Gatekeeper technology work
Gatekeeper is a multi-factor voice authentication solution that uses audio signal containing speech as its primary input. It extracts Features from the audio signal using Mel-frequency Cepstral Coefficients (MFCCs), passes them to a series of neural networks, and generates outputs and scores.
The four main output scores consist of Verification, Synthetic Speech Detection, Channel Playback, and Known Fraudster Detection. The combination of these factors generates a combined Session Risk Score, which is consumed by customer applications and contact center agents as an unambiguous indication of risk.
- Verification: 1:1 comparison of previously recorded speaker and current speaker
- Synthetic Speech Detection: Recognition of non-natural speech patterns and artifacts of digitally manipulated speech
- Playback Detection: Recognition of loudspeakers
- Known Fraudster Detection: 1:n comparison against a list of customer labeled fraudster voices
Audio Signal Input

Neural Network Processing

Customer consumed output

How is performance measured
Biometric voice authentication technology is robust and probabilistic in nature: Incorrect decisions are normal and to be expected. Different types of decisioning errors occur and customer selected operating points determines the proportion of these errors. There’s a delicate balance between the proportion of correct and incorrect decisions that are measured as follows:
- False Accept Rate (FA): % of imposters that receive a positive verification decision
- False Reject Rate (FR): % of true users that receive a negative verification decision
The relationship between the two are inversely related, the higher the security setting (lower false accept rate) implies more legitimate customers failing authentication (higher false reject rate).

What affects performance
The following factors are considered when assessing performance:
-
Amount of audio available. Typical uses cases vary between 2—15 seconds of net speech.
-
Verification versus non-verification audio. Is the end user enrolled in voice biometrics? Voiceprints are unique to each Gatekeeper customer and an identity claim (for example, account number) must be made to attempt verification.
The relative performance of FA/FR trade-off can be resumed in the table below:

-
Engine version. The more recent the version, the higher the overall performance. Version 10 and 11 can all achieve synthetic speech detection rates of 90%+ with false reject rates varying between 2% to 10%. Newer engine = same security, lower customer impact.
-
Engine calibration. Gatekeeper is delivered with factory models. These are pre-calibrated based upon target FA/FR rate using internal training data sets. Factory models are rarely used due to high variability of customer audio quality and security settings (customer convenience versus optimal fraud detection versus high sensitivity to synthetic speech). Professional services custom calibrates and deploys Gatekeeper engines using customer production data as a training input. Custom trained models improve the equal error rate of factory version 11 model by approximately 20%.
-
Audio quality. Use case, customer telephony platform configuration, network carrier, type of business, geographic region, customer devices, and calling behavior can all have material impacts on voice authentication performance.
High security can be achieved with relatively low FR when a typical customer calls their investment bank to transfer a $1M. They make their calls from a quiet office using a high-end smartphone with a strong voice-over wi-fi connection. Alternatively, a customer calling to negotiate their prepaid cell phone provider while in the car in a rural area, on speakerphone has significant higher FR for the same FA due to the lower quality audio signal input used to render a decision.
Customer-specific configuration and calibration
Gatekeeper platform supports a wide range of use cases ranging from voice authentication in a retail store, authenticating high risk investment banking transactions, screening of new account openings as well as day to day customer interactions through the telephony channel for banks, telcos, IT help desks and retailers alike.
Due to the variability in use cases and customer-specific security versus convenience tolerance, Gatekeeper performance is highly customizable. How much and what type of audio a decision is rendered with, what False Accept rate is targeted, sensitivity of playback and synthetic speech attacks, and how the outputs are integrated into both customer applications and business processes all have dramatic effects on performance. Here are two opposing examples:
-
IT help desk use case:
- 1% False Accept target
- 50% synthetic speech detection target
- 50% playback detection target
- 0% known fraudster detection target
- 1% False Reject Rate
-
Investment banking transaction use case for live agents:
- 0.1% False Accept target
- 99% synthetic speech detection target
- 95% playback detection target
- 95% known fraudster detection target
- 8% False Reject Rate
Customer often implement custom integrations, custom call flows, custom engine calibrations, and thresholds established by professional service providers through consultation with customer stakeholders.
Generic product-level performance metrics can be misleading and are meaningful only when standard product configurations and calibrations are used.
Such flexibility also allows customers to have multitude of Risk Engine calibrations on hand to be able to adapt the security versus convenience settings in an autonomous and agile way as threats or operational realities change.
Synthetic speech detection performance
There is no known definition or standard defining synthetic speech or detection performance within contact center environments. Current publicly available benchmarks are based on digital content and audiobook reading. They neither currently represent real world natural speech present in contact centers, nor the rapid advancements in generative voice AI.
It must also be stated that synthetic speech isn’t inherently nefarious and significant amount of non-natural speech is present in public telephony and contact centre environments. Low quality Bluetooth devices and speakerphone can introduce crossover between the speaker and microphone. Low quality VOIP networks can compress and denature the audio signal. Hold music, automated welcome messages, IVR prompts, voice auto attendants, and background music are all examples of legitimate synthesized audio that are common in real-world environments.
These types of audio aren’t by definition ‘false alerts.’ In terms of synthetic speech detection, they’re true-positives as they correctly detect and identify non-natural speech.
Common perception is that these ‘false alerts’ are a technological deficiency. To be able to render decisions that provide the highest level of protection against voice deepfakes, customers can increase their minimum acceptable standards of audio quality to securely service their customers.
Channel Playback and Synthetic Speech Detection factors are always used together to provide protection against a broad set of attack vectors. Depending on how the synthetic audio is presented to the Gatekeeper system, different factors are leveraged to detect the attack.

Model training and data sets
Gatekeeper technology is developed using public, private, and proprietary data sets. With the emergence and rapid growth of generative voice technology, no benchmark data or testing methodology currently exists to assess or quantify performance of voice authentication technology. Millions of proprietary samples of synthetic audio are generated to emulate third-party fraudster synthetic speech attacks.
The third-party fraud attack vector is another careful consideration when building and training data sets. Seeing as this threat has yet to materialize, Gatekeeper aims to provide broad protection against a multitude of synthetic speech attacks that may occur in the future.
How much victim audio is available for synthesis purposes? What software is used to synthesize the audio? How much synthetic audio is presented for verification? What type of device is used to present the audio? These are just a subset of variables used when building our technology. Over 1,500 different attack scenarios are modeled.
The following assumptions are made in regard to model development, training and testing:
- The fraudster can’t access the Gatekeeper Admin Console or API interface.
- Fraud attack presents the audio through the same channel that’s used in production (for example, synthesized audio is presented to an IVR or contact center channel through the public telephony network).
- Victim audio is acquired covertly by the fraudster.
- Victim doesn’t voluntarily provide audio to the fraudster. For example, the victim’s audio is provided to synthetic speech software directly from their smartphone microphone or professional sound recording equipment.
- Covert victim audio acquisition occurs from a probable source (for example, use of social media audio, covert recording of a phone conversation through a telephony channel).
- The fraudster doesn’t have physical access to the victim’s device or environment.
- It’s highly improbable that a third-party fraudster uses the victim’s device to directly record their voice for synthesis purposes or to present synthesized audio to Gatekeeper from the same environment as the victim.
Synthetic speech as a fraud detection input
Across our global customer base with billions of voice authenticated calls, synthetic speech threat, while real, has yet to materialize. Fraudsters overwhelmingly attack victim accounts that don’t possess a voiceprint.
Synthetic speech and recorded audio are present in contact center and telephony audio for legitimate and non-fraudulent reasons. Hold music, background music, voice auto attendants, agent welcome message, voicemail greeting messages, and garbled audio are just some examples of non-natural speech patterns that occur in production environments.
Using Gatekeeper synthetic speech capabilities as an input for customer fraud operations teams may result in many alerts with virtually zero true fraud alerts. This paradigm may shift in the future, but currently synthetic speech and playback detection should be seen as a safety net that protect against future threats. The operational impact of this safety net must be carefully assessed and configured by each customer.
Multiple Risk Engine calibrations of varying sensitivities can be stored in a Gatekeeper scope to be able to adapt the security versus convenience settings in an autonomous and agile way as threats or operational realities change.
Synthetic speech performance versus penetration type tests
With the absence of industry and third-party benchmark testing, evaluating synthetic speech and playback detection performance trade-offs are challenging for some customers to assess. Customer information security teams are sometimes involved to assess Gatekeeper performance using cyber security ‘penetration testing’ approaches. Gatekeeper voice authentication performance or synthetic speech detection performance doesn’t fall under formal Microsoft penetration testing considerations.
We welcome customers and support them when executing their own testing strategies. There are several aspects that differ from unauthorized access or information leak type penetration tests:
- Synthetic speech and playback audio are naturally occurring events in production systems that aren’t nefarious in nature and don’t equate to unauthorized access.
- False accepts are normal and expected for any probabilistic decisioning system.
- Decisioning is probabilistic in nature, tests executed under ideal circumstances (first-party attack) are likely to generate a False Accept. While Gatekeeper is able to detect such events with stricter configuration settings, they aren’t representative of third-party fraud attacks in production systems.
- Creating statistically representative or meaningful test cases involve recruiting many individuals to provide their voices to third-party synthetic speech providers. A minimum of 25 different individuals is recommended.
- Recreating and executing a variety of probable third-party attack vectors can take significant resources and materials.
- Being overly sensitive to synthetic speech and playback audio can increase False Reject rates and use of higher risk fallback authentication methods.
- Using either a small number of testers or ideal testing scenarios can lead to overfitting of calibrations/thresholds to risks that are unrealistic fraud threats.