WebRTC

WebRTC SDP Explained: How the Offer/Answer Model Works, Line by Line

18 min read
WebRTC
Reading Time: 12 minutes

Every WebRTC call starts with two peers swapping a block of text. That text is a WebRTC SDP, and if it’s wrong, nothing else matters. No audio, no video, no data channel.

SDP (Session Description Protocol) is how two browsers agree on codecs, encryption keys, network addresses, and which media streams they’ll send. It’s also the part of WebRTC most developers copy from a tutorial and never read.

That works right up until a call fails with Failed to set remote answer sdp and you’re staring at 80 lines of a= attributes.

Once you can read an SDP, most of those failures explain themselves. Below you’ll find the offer/answer exchange in code, a real WebRTC SDP decoded line by line, and a table of the errors you’ll hit most often.

What Is SDP in WebRTC?

SDP (Session Description Protocol) is a text-based format that describes a multimedia session: which media types will flow, which codecs each side supports, how the media is encrypted, and how to reach each peer on the network. WebRTC uses SDP as the payload of its offer/answer negotiation, which the browser generates through the RTCPeerConnection API.

SDP isn’t a transport protocol. It doesn’t move audio or video. It’s a description that both sides read before any media flows, closer to a contract than a conversation.

The format predates WebRTC by more than a decade. It was built for SIP phone calls and streaming.

The current spec is RFC 8866, published in January 2021 to replace RFC 4566. WebRTC adds its own rules on top through JSEP (JavaScript Session Establishment Protocol), which defines exactly which SDP lines a browser produces and accepts.

Property What SDP does in WebRTC
Format Plain text, one key=value pair per line
Keys Single letters (v, o, s, t, c, m, a)
Carried by Your own signaling channel (WebSocket, HTTP, etc.)
Created by createOffer() and createAnswer()
Applied with setLocalDescription() and setRemoteDescription()
Describes Media sections, codec information, ICE credentials, DTLS fingerprint, stream IDs
Doesn’t do Transport media, traverse NAT, or encrypt anything by itself

SDP meaning outside WebRTC

If you search “what is SDP,” you’ll find the acronym used for software-defined perimeters, security products, and even payroll systems. In real-time video, SDP always means Session Description Protocol.

You’ll also see .sdp files used with RTSP and FFmpeg to describe an RTP stream. An SDP file is the same format saved to disk instead of sent over a signaling channel.

Where SDP Fits in a WebRTC Connection

SDP is one piece of a larger handshake, and it feeds every other piece.

A WebRTC connection needs four things to happen:

  1. Signaling: Peers exchange SDP offers, answers, and ICE candidates through a server you run. WebRTC deliberately doesn’t define this channel.
  2. Connectivity (ICE): Interactive Connectivity Establishment tests candidate address pairs to find a path between peers, using STUN and TURN when needed.
  3. Security (DTLS): The peers run a DTLS handshake and check the certificate against the fingerprint in the SDP. SRTP keys come from that handshake.
  4. Media (RTP/SRTP): Encrypted audio and video flow using the codecs and payload types agreed in the SDP.

It carries the ICE username and password, the DTLS fingerprint and role, and the RTP payload mapping. If one of those values is wrong, the step that depends on it fails, often with an error that doesn’t mention SDP at all.

That’s also why you need a WebRTC signaling server. The browser builds the SDP, but it has no built-in way to deliver it to the remote peer. You pick the transport: WebSocket, HTTP POST, even copy and paste for a demo.

How the SDP Offer/Answer Model Works

WebRTC negotiates with the offer/answer model defined in RFC 3264. One peer proposes a session. The other accepts what it can and rejects the rest.

The exchange is asymmetric: the answer can only narrow what the offer proposed, never add to it.

Here’s the sequence between two peers, Alice (the caller) and Bob:

  1. Alice creates an SDP offer with createOffer(). It lists every codec and media section she’s willing to use.
  2. Alice calls setLocalDescription(offer). This locks in her local description and starts ICE candidate gathering.
  3. Alice sends the offer message to Bob through the signaling server.
  4. Bob calls setRemoteDescription(offer). His browser now knows what Alice supports.
  5. Bob creates an SDP answer with createAnswer(). It keeps only the codecs and sections he also supports.
  6. Bob calls setLocalDescription(answer) and sends the answer back.
  7. Alice calls setRemoteDescription(answer). Negotiation is done. ICE checks and the DTLS handshake follow.

In code, the caller side looks like this:

const pc = new RTCPeerConnection({
  iceServers: [{ urls: 'stun:stun.l.google.com:19302' }]
});

// Add local tracks so they appear as m= sections in the offer
const stream = await navigator.mediaDevices.getUserMedia({ audio: true, video: true });
stream.getTracks().forEach(track => pc.addTrack(track, stream));

// Send ICE candidates as they're found (trickle ICE)
pc.onicecandidate = ({ candidate }) => {
  if (candidate) signaling.send({ type: 'candidate', candidate });
};

// Create and send the offer
const offer = await pc.createOffer();
await pc.setLocalDescription(offer);
signaling.send({ type: 'offer', sdp: pc.localDescription.sdp });

// Apply the answer when it arrives
signaling.on('answer', async ({ sdp }) => {
  await pc.setRemoteDescription({ type: 'answer', sdp });
});

And the callee side:

signaling.on('offer', async ({ sdp }) => {
  await pc.setRemoteDescription({ type: 'offer', sdp });

  const answer = await pc.createAnswer();
  await pc.setLocalDescription(answer);
  signaling.send({ type: 'answer', sdp: pc.localDescription.sdp });
});

signaling.on('candidate', async ({ candidate }) => {
  await pc.addIceCandidate(candidate);
});

Order matters. Call setRemoteDescription() before createAnswer(), and add remote candidates only after a remote description exists. Break either rule and you’ll get an InvalidStateError.

SDP types beyond offer and answer

RTCSessionDescription has four possible type values:

  • offer: A proposal to start or change a session.
  • answer: The final response that completes negotiation.
  • pranswer: A provisional answer. It’s mostly used when bridging to SIP systems and rarely appears in browser-to-browser apps.
  • rollback: Cancels a pending offer and returns to the last stable state. It’s the key to handling “glare,” where both peers send an offer at the same moment.

Anatomy of a WebRTC SDP

Here’s a trimmed SDP offer from Chrome for one audio track and one video track. Real offers run 80 to 150 lines, but the structure is the same.

v=0
o=- 4611731400430051336 2 IN IP4 127.0.0.1
s=-
t=0 0
a=group:BUNDLE 0 1
a=extmap-allow-mixed
a=msid-semantic: WMS
m=audio 9 UDP/TLS/RTP/SAVPF 111 63 9 0 8 13 110 126
c=IN IP4 0.0.0.0
a=rtcp:9 IN IP4 0.0.0.0
a=ice-ufrag:Kq2x
a=ice-pwd:3Gm8pX0vT5bW9sR1yL7nQe2A
a=ice-options:trickle
a=fingerprint:sha-256 7B:8B:F0:65:5F:78:E2:51:3B:AC:6F:F3:3F:46:1B:35:DC:B8:5F:64:1A:24:C2:43:F0:A1:58:D0:A1:2C:19:08
a=setup:actpass
a=mid:0
a=extmap:1 urn:ietf:params:rtp-hdrext:ssrc-audio-level
a=sendrecv
a=msid:stream1 audiotrack1
a=rtcp-mux
a=rtpmap:111 opus/48000/2
a=rtcp-fb:111 transport-cc
a=fmtp:111 minptime=10;useinbandfec=1
m=video 9 UDP/TLS/RTP/SAVPF 96 97 98 99 45
c=IN IP4 0.0.0.0
a=rtcp:9 IN IP4 0.0.0.0
a=ice-ufrag:Kq2x
a=ice-pwd:3Gm8pX0vT5bW9sR1yL7nQe2A
a=ice-options:trickle
a=fingerprint:sha-256 7B:8B:F0:65:5F:78:E2:51:3B:AC:6F:F3:3F:46:1B:35:DC:B8:5F:64:1A:24:C2:43:F0:A1:58:D0:A1:2C:19:08
a=setup:actpass
a=mid:1
a=sendrecv
a=msid:stream1 videotrack1
a=rtcp-mux
a=rtcp-rsize
a=rtpmap:96 VP8/90000
a=rtcp-fb:96 nack
a=rtcp-fb:96 nack pli
a=rtcp-fb:96 transport-cc
a=rtpmap:97 rtx/90000
a=fmtp:97 apt=96
a=rtpmap:98 H264/90000
a=fmtp:98 level-asymmetry-allowed=1;packetization-mode=1;profile-level-id=42e01f
a=rtpmap:99 rtx/90000
a=fmtp:99 apt=98
a=rtpmap:45 AV1/90000

An SDP has two layers. The session section comes first and applies to everything.

Each media section starts with an m= line and runs until the next m= line. Attributes inside a media section apply only to that stream.

Session-level lines

These first lines are mostly fixed values in WebRTC. JSEP tells browsers exactly what to put in them.

Line Example Meaning
v= v=0 Protocol version. Always 0.
o= o=- 4611731400430051336 2 IN IP4 127.0.0.1 Origin: username (-), session ID, session version, network type, address. The version number goes up on every renegotiation.
s= s=- Session name. WebRTC leaves it as a dash.
t= t=0 0 Start and stop time. 0 0 means the session is permanent.
a=group:BUNDLE a=group:BUNDLE 0 1 Sends all listed media sections over one transport.
a=msid-semantic a=msid-semantic: WMS Declares that msid lines map tracks to MediaStreams.

The address in the o= line (127.0.0.1) is a placeholder. It’s never used for routing, so don’t worry that it isn’t your real public IP address.

The m= line: media type, transport, and payload types

m=video 9 UDP/TLS/RTP/SAVPF 96 97 98 99 45

Read it left to right:

  • video: The media type. Others are audio and application (used for the data channel).
  • 9: The port. It’s a dummy value because ICE candidates carry the real ports.
  • UDP/TLS/RTP/SAVPF: The transport profile. It means RTP over DTLS-SRTP over UDP, with RTCP feedback. A WebRTC data channel section uses UDP/DTLS/SCTP instead.
  • 96 97 98 99 45: The RTP payload types offered, in order of preference. Each number is defined by an a=rtpmap line further down.

ICE attributes

a=ice-ufrag:Kq2x
a=ice-pwd:3Gm8pX0vT5bW9sR1yL7nQe2A
a=ice-options:trickle

ice-ufrag and ice-pwd are short-lived credentials. The remote peer uses them to sign STUN connectivity checks, so each side can confirm it’s talking to the peer it negotiated with. ice-options:trickle says this peer will send ICE candidates as it finds them, instead of waiting until gathering finishes.

When candidates do appear in the SDP (in non-trickle mode, or inside localDescription after gathering), they look like this:

a=candidate:1 1 udp 2122260223 192.168.1.20 54400 typ host
a=candidate:2 1 udp 1686052607 203.0.113.7 54400 typ srflx raddr 192.168.1.20 rport 54400
a=candidate:3 1 udp 41885439 198.51.100.4 3478 typ relay raddr 203.0.113.7 rport 54400

host is a local address. srflx (server reflexive) is your public address as seen by a STUN server. relay is an address on a TURN server that forwards traffic when a direct path isn’t possible. Our guide to NAT traversal covers how ICE picks between them.

DTLS attributes: fingerprint and setup

a=fingerprint:sha-256 7B:8B:F0:65:...
a=setup:actpass

These are the security parameters. The fingerprint is a hash of the certificate the peer will present during the DTLS handshake, and a mismatch drops the connection.

That’s how WebRTC blocks a man-in-the-middle from swapping in its own keys, as long as your signaling channel is itself secure.

a=setup decides who acts as the DTLS client:

  • actpass: “I can be either.” Offers always use this.
  • active: “I’ll be the client.” Most answers use this.
  • passive: “I’ll be the server.”

Media identity: mid and msid

a=mid:0 gives each media section an ID. BUNDLE references these IDs, and so does RTP through a header extension, so a receiver knows which section an incoming packet belongs to.

a=msid:stream1 videotrack1 ties the section to a MediaStream ID and a track ID. That’s what lets the remote ontrack event hand you the right stream.

a=sendrecv sets direction. The other values are sendonly, recvonly, and inactive. A viewer who only watches a stream answers with recvonly.

Codec lines: rtpmap, fmtp, and rtcp-fb

This is where most of the codec information lives.

a=rtpmap:111 opus/48000/2
a=fmtp:111 minptime=10;useinbandfec=1
a=rtpmap:96 VP8/90000
a=rtcp-fb:96 nack pli
a=rtpmap:97 rtx/90000
a=fmtp:97 apt=96
  • rtpmap maps a dynamic RTP payload type number to a codec name, clock rate, and (for audio) channel count. Payload type 111 is the Opus codec at 48 kHz stereo. Video codecs use a 90 kHz clock.
  • fmtp carries codec-specific settings. For Opus, useinbandfec=1 turns on forward error correction. For H.264, profile-level-id=42e01f means Constrained Baseline, level 3.1.
  • rtcp-fb lists the feedback each codec supports. nack requests retransmission of lost packets. nack pli asks for a new keyframe. transport-cc enables transport-wide congestion control.
  • rtx is a retransmission stream. apt=96 says it carries resent packets for payload type 96.

Don’t hard-code payload type numbers. The 96 to 127 range is dynamic, and Chrome and Firefox assign different numbers to the same codec, and the answer must reuse whatever numbers the offer chose.

WebRTC browsers must support VP8 and H.264 for video, plus Opus and G.711 for audio. Most now offer VP9 and AV1 as well. Our guides to video codecs and the VP9 codec cover the tradeoffs between them.

RTP header extensions and RTCP

a=extmap:1 urn:ietf:params:rtp-hdrext:ssrc-audio-level
a=rtcp-mux
a=rtcp-rsize

extmap lines negotiate RTP header extensions, which attach small bits of metadata to each packet. Audio level, absolute send time, video orientation, and the mid value all travel this way. rtcp-mux sends RTCP on the same port as RTP, and rtcp-rsize allows smaller RTCP packets.

Simulcast lines

When a sender encodes the same video at several resolutions, the SDP includes lines like these:

a=rid:h send
a=rid:m send
a=rid:l send
a=simulcast:send h;m;l

Each rid names one encoding layer, so an SFU can forward the high layer to viewers on fiber and the low layer to viewers on mobile.

Simulcast in SDP is defined in RFC 8853. Scalable video coding (SVC) with VP9 or AV1 does something similar inside a single stream, and it’s signaled through scalabilityMode in the API instead.

Unified Plan vs Plan B

Older articles and Stack Overflow answers often show SDPs with one m=video section containing several a=ssrc groups, one per track. That’s Plan B, a Chrome-only format.

The standard is Unified Plan: one m= section per track. Three video tracks mean three m=video sections.

Chrome switched its default to Unified Plan in 2019 and has since removed Plan B entirely, so any SDP parsing code written for Plan B needs rewriting.

Plan B Unified Plan
Tracks per m= section Many One
Track identity a=ssrc lines a=mid and a=msid
Browser support Removed from Chrome All modern browsers
API model Streams Transceivers

Trickle ICE and Renegotiation

Two behaviors change what an SDP looks like over the life of a call.

Trickle ICE. Early WebRTC apps waited for all ICE candidates before sending the offer, which could add several seconds when a TURN server was slow to respond.

With trickle ICE, you send the SDP right away and stream candidates separately through onicecandidate. It’s the default in every modern browser and cuts call setup time.

Renegotiation. A new screen share track, a muted video, or a camera switch can all change the session. When that happens, the browser fires negotiationneeded, and you run a new offer/answer round.

The o= line’s version number goes up, while the session ID stays the same. Tutorials on WebRTC screen sharing often trip over this step because a new track adds a new m= section.

To avoid glare when both peers renegotiate at once, use the “perfect negotiation” pattern: one peer is “polite” and rolls back its own offer when a collision happens. The other is “impolite” and ignores incoming offers during its own.

SDP Munging: Editing the SDP by Hand

SDP munging means changing the SDP string between createOffer() and setLocalDescription(). Developers used to munge to force a codec, cap bitrate with a b=AS: line, or turn on stereo Opus.

It still works in some cases, but it’s fragile. The browser may reject edits it doesn’t expect, and the spec now discourages munging the local description.

Use the API instead when you can:

  • Codec preference: RTCRtpTransceiver.setCodecPreferences()
  • Bitrate caps: RTCRtpSender.setParameters() with encodings[].maxBitrate
  • Simulcast layers: sendEncodings in addTransceiver()
  • Direction: transceiver.direction = 'recvonly'

If you must munge, edit only fmtp values or remove codecs you don’t want. Never touch ICE credentials, fingerprints, or mid values.

Common WebRTC SDP Errors and How to Fix Them

Most SDP failures come from a handful of causes. Here’s what each console error actually means.

Error or symptom Likely cause Fix
InvalidStateError: Called in wrong state: stable Setting an answer when no offer is pending, or processing a duplicate answer Check pc.signalingState before applying; dedupe messages
Failed to set remote answer sdp: The order of m-lines in answer doesn't match order in offer The answer was edited, or built from a different offer Never reorder m= sections; answer the exact offer you received
Failed to set remote offer sdp: Session error code: ERROR_CONTENT No codec in common, or malformed rtpmap/fmtp after munging Remove munging; confirm both sides support at least one codec
Connection stuck in checking, then failed ICE can’t find a path: no TURN server, or candidates dropped Add TURN; make sure every candidate reaches the other peer
addIceCandidate throws before connection starts Candidate arrived before the remote description was set Queue candidates until setRemoteDescription() resolves
Video negotiated but black screen on one side Direction mismatch (recvonly/inactive) or no track added Check a=sendrecv values and that addTrack() ran before the offer
DTLS fails right after ICE connects Fingerprint altered in transit, or setup roles both active Don’t modify fingerprints; let the answer choose active or passive

To debug, log both SDPs as plain text on each side and compare them.

In Chrome, chrome://webrtc-internals shows every offer, answer, and candidate with timestamps. Firefox has the same data at about:webrtc.

From Peer-to-Peer SDP to Streaming at Scale

Everything above describes a two-party call. Past a handful of participants, the SDP picture changes, and so does the infrastructure you have to run.

In a group call, each browser negotiates one SDP with a media server instead of with every other browser. That server is usually an SFU (Selective Forwarding Unit), which receives each sender’s stream once and forwards it to everyone else.

The server now owns half of every offer/answer exchange. It has to manage codec negotiation, simulcast layers, and renegotiation as people join and leave. Our overview of WebRTC servers compares the architectures.

SDP in WHIP and WHEP

Broadcasting over WebRTC has its own standard now. WHIP (WebRTC-HTTP Ingestion Protocol), published as RFC 9725 in March 2025, replaces the custom signaling server with one HTTP request:

curl -X POST https://media.example.com/whip/endpoint \
  -H "Content-Type: application/sdp" \
  -H "Authorization: Bearer $TOKEN" \
  --data-binary @offer.sdp

The encoder POSTs its SDP offer, and the server replies with the SDP answer. WHEP does the same for playback.

It’s the simplest way to see that SDP is just text in the body of a request.

When WebRTC isn’t the right delivery layer

WebRTC gets glass-to-glass latency under 500 ms, which is what you want for calls, auctions, and interactive shows.

But every viewer needs a live peer connection, its own SDP negotiation, and possibly a TURN relay. At 10,000 viewers, that’s 10,000 stateful sessions to run.

For broadcasts to large audiences, many teams send the stream in over RTMP or SRT and deliver it as HLS through a CDN instead. Latency rises to a few seconds, but the delivery side scales on commodity HTTP caching. Our comparison of WebRTC vs HLS breaks down where each one fits, and the guide to ultra-low latency streaming covers options in between.

This is where LiveAPI fits. You don’t write any SDP. You push a stream over RTMP or SRT, and LiveAPI handles:

  • Transcoding to adaptive bitrate HLS
  • Delivery across Akamai, Cloudflare, and Fastly
  • Recording every stream to VOD automatically

You can mix the two models too: run WebRTC for the handful of people on stage, then hand the composed output to LiveAPI for the audience. Our guide to WebRTC live streaming shows how that hybrid setup works.

WebRTC SDP FAQ

What is SDP in WebRTC?

SDP is the text format WebRTC peers use to describe their session before connecting. It lists media tracks, supported codecs, ICE credentials, and the DTLS certificate fingerprint. Peers exchange SDP through a signaling channel as an offer and an answer.

What is the SDP protocol?

The Session Description Protocol is an IETF standard, currently RFC 8866, for describing multimedia sessions. It defines a line-based key=value format but doesn’t transport any media. SIP, RTSP, and WebRTC all use it to describe sessions.

What is an SDP file?

An SDP file is a session description saved to disk with a .sdp extension. Tools like FFmpeg and VLC read SDP files to receive raw RTP streams, since the file tells them the codec, payload type, and port. In WebRTC, the same content usually lives in memory and travels over signaling instead.

Is SDP encrypted?

No. SDP is plain text, and WebRTC doesn’t encrypt it. The media is encrypted with DTLS-SRTP, but the fingerprint that protects that handshake travels inside the SDP. That’s why your signaling channel should always use HTTPS or WSS.

What’s the difference between the local description and the remote description?

The local description is the SDP your peer generated and applied with setLocalDescription(). The remote description is the SDP you received from the remote peer and applied with setRemoteDescription(). You need both before media can flow.

Does WebRTC require a signaling server to exchange SDP?

Yes, in practice. WebRTC generates the SDP but doesn’t say how to deliver it. Most apps use WebSockets, though WHIP and WHEP use plain HTTP, and simple demos can exchange SDP through copy and paste.

Can I change codecs after the call starts?

Yes. Call setCodecPreferences() on the transceiver and then run a new offer/answer round. The browser fires negotiationneeded when a change requires it, and the new SDP carries the updated codec order.

Why does the SDP show 0.0.0.0 and port 9?

They’re placeholders. Real addresses and ports come from ICE candidates, which are exchanged separately with trickle ICE. JSEP tells browsers to fill c= and m= lines with these dummy values.

Put Your WebRTC SDP Knowledge to Work

A WebRTC SDP looks dense, but it follows a fixed pattern: a session section, one media section per track, and a set of attributes for ICE, DTLS, and codecs.

Once you can read those groups, most negotiation errors point straight to their cause.

For calls and small interactive sessions, owning that negotiation makes sense. For broadcasts to thousands of viewers, you can skip it and send your stream through an API built for delivery at scale. Get started with LiveAPI to go live over RTMP or SRT and reach viewers everywhere through HLS.

Join 200,000+ satisfied streamers

Still on the fence? Take a sneak peek and see what you can do with Castr.

No Castr Branding

No Castr Branding

We do not include our branding on your videos.

No Commitment

No Commitment

No contracts. Cancel or change your plans anytime.

24/7 Support

24/7 Support

Highly skilled in-house engineers ready to help.

  • Check Free 7-day trial
  • CheckCancel anytime
  • CheckNo credit card required

Related Articles