The message
You receive ananswering_machine message at most once per session:
time is measured on the audio timeline, not the clock. It uses the same reference as speech_start, speech_end and transcript, so you can order all of them against each other. If you push audio faster than real time, time still refers to the audio position, so use created_at when you need to correlate with your own logs.Acting on it
Most callers act onkind directly:
machine— hang up and queue the number for a later attempt, or stay silent and let the greeting finish before leaving a recorded message.human— start the conversation.
confidence and require a high score before treating a call as a machine. Hanging up on a real person usually costs more than a few wasted seconds on a voicemail, so a threshold above the default is a common choice.
Timing and edge cases
The decision fires at most once per session and is never revised. Later audio does not change it.- The window starts at the first speech. Ringing, connection delay and any leading silence do not count. The decision is taken once five seconds of audio have been heard after that first speech.
- Nobody speaks at all. No message is emitted: silence never starts the window and is not evidence of a machine. Apply your own timeout if you need to give up on a silent line.
- Very short calls. If the session ends before the five seconds have elapsed, no message is emitted.
- Transcripts are not ordered against it. Partial or final transcripts may arrive before the decision.
- No transcription required. The detection does not depend on what was said, so it behaves the same across languages and adds no transcription latency.