One browser method does all of it. navigator.mediaDevices.getDisplayMedia() is the only supported way for a web page to capture a screen, a window, or a tab. It ships in every major desktop browser, and it’s missing from every browser on iOS.
That single fact shapes most of the screen sharing architecture decisions you’ll make.
WebRTC screen sharing hands your app a live MediaStream of someone’s display, then moves it peer to peer with the same encryption, congestion control, and sub-second delivery it uses for camera video.
What Is WebRTC Screen Sharing?
WebRTC screen sharing is a browser capability that captures the contents of a display surface (an entire monitor, an application window, or a single browser tab) as a real-time MediaStream and transmits it to remote peers over an encrypted WebRTC peer connection.
Three things define it.
The capture side runs through the Screen Capture API’s getDisplayMedia() method, which always requires an explicit user gesture and a browser-rendered picker. The transport side runs through RTCPeerConnection, the same stack that carries camera and microphone media, so the shared screen inherits DTLS-SRTP encryption, adaptive bitrate behavior, and packet loss recovery.
And the whole thing runs without a plugin or an extension, which wasn’t true before 2018.
Screen sharing exists as a separate API from camera capture for a privacy reason. A page that can silently read your display can read your password manager, your email, and your bank tab.
So the spec forces the user agent to own the picker, blocks source enumeration, and shows a system-level indicator while capture is active. Your JavaScript can’t route around any of it.
getDisplayMedia vs getUserMedia: What’s the Difference?
Developers new to WebRTC often reach for getUserMedia() and wonder why there’s no “screen” device in the list. There isn’t one, by design.
getUserMedia() captures cameras and microphones: hardware devices that appear in enumerateDevices(). getDisplayMedia() captures display surfaces, which never appear in device enumeration and never fire devicechange events.
Both return a MediaStream, and both feed the same peer connection. The permission models are where they split.
| Aspect | getUserMedia() | getDisplayMedia() |
|---|---|---|
| Captures | Camera, microphone | Monitor, window, browser tab |
Source listed in enumerateDevices() |
Yes | No |
| Permission model | Persistent, per-origin grant | Per-call, user picks the target every time |
| Requires a user gesture | Not always | Yes |
| Permissions Policy directive | camera, microphone |
display-capture |
| Audio track | Reliable | Optional, platform-dependent |
| Mobile browser support | Broad | Desktop only |
The per-call permission model matters more than it looks. You can’t cache a screen share grant the way you cache camera access, so every new share means a new picker dialog. Build your UI around that instead of fighting it.
How Does WebRTC Screen Sharing Work?
The pipeline from click to remote render has seven steps. Most are identical to a camera call. Only the first two are specific to screen capture.
- A user gesture triggers capture. A click handler calls
getDisplayMedia(). The browser renders its own picker listing tabs, windows, and screens. Your page never sees the list. - The browser returns a MediaStream. After the user picks a surface, you get a stream with one video track and, sometimes, one audio track. The track carries the real resolution and frame rate the OS is producing, which may not match what you asked for.
- The track is added to a peer connection.
addTrack()attaches it to an existingRTCPeerConnection, orreplaceTrack()swaps it into a sender that’s already carrying camera video. - Signaling exchanges the new SDP. Adding or replacing a track fires
negotiationneeded. Your signaling server relays the new offer and answer between peers over whatever channel you chose: WebSocket, SSE, or a message queue. - ICE finds a path. Candidates gathered from a STUN server let peers learn their public addresses and connect directly. When symmetric NAT or a strict firewall blocks that, media relays through a TURN server instead.
- Media flows over SRTP. The encoder compresses frames, congestion control adjusts the bitrate to the measured path capacity, and the receiver’s jitter buffer smooths out arrival timing.
- The remote peer renders the stream. The receiving app sets the incoming stream as a video element’s
srcObject, and frames appear typically within 100–300 ms of being drawn on the sender’s screen.
The step that surprises people is the fourth one. Adding a screen track to a live call renegotiates the session, so a screen share that works on localhost can fail in production when your signaling layer drops or reorders that second offer.
Types of WebRTC Screen Capture: Tab, Window, and Screen
The user picks what gets shared, but your displaySurface hint decides which tab the picker opens on and which panes it shows. Each surface type behaves differently once capture starts.
1. Browser Tab Capture
Setting displaySurface: "browser" opens the picker on the tab list.
Tab capture is the cleanest option: resolution follows the tab’s viewport, the frame rate stays stable, and it’s the only surface type where audio capture works reliably across Chrome, Edge, and Opera. It also survives the user switching windows, since the tab keeps rendering in the background.
Best for: web apps demoing other web apps, support agents walking through a dashboard, and any share where tab audio matters.
2. Application Window Capture
displaySurface: "window" shows the list of open application windows. The capture follows that one window, so nothing else on the desktop leaks into the stream.
The tradeoff is that content the OS never composites (a window sitting behind another window on some platforms) may come through blurred or stale, because the browser is capturing a logical surface rather than what’s literally on the glass.
Best for: sharing a native app, an IDE, or a design tool, where privacy for the rest of the desktop is the priority.
3. Entire Screen (Monitor) Capture
displaySurface: "monitor" captures a whole display, notifications and all. It’s the highest-risk surface from a privacy standpoint and the most bandwidth-hungry, since a 4K monitor produces far more pixels than a tab viewport. Chrome 119+ lets you remove the option entirely with monitorTypeSurfaces: "exclude".
Best for: presentations that span several applications, remote support, and training sessions.
4. Screen Audio
Audio is the least portable part of screen capture. Tab audio works on Chrome, Edge, and Opera. System-wide audio capture works on Windows and, in a limited way, on ChromeOS. It isn’t available on macOS or Linux, and Safari doesn’t offer screen audio at all.
Treat it as a bonus feature and always give users a microphone fallback.
| Surface type | Constraint value | Audio capture | Bandwidth | Privacy exposure |
|---|---|---|---|---|
| Browser tab | "browser" |
Reliable on Chromium | Lowest | Lowest |
| Application window | "window" |
Rare | Medium | Low |
| Entire screen | "monitor" |
Windows only | Highest | Highest |
WebRTC Screen Sharing Architectures: P2P, SFU, and MCU
Capture is the easy half. How the stream reaches viewers decides whether your app works with three people or three hundred.
In a peer-to-peer mesh, the sharer opens a direct connection to every viewer and encodes the screen once per connection. At roughly 1.5 Mbps for a legible 1080p share, a presenter with five viewers is pushing 7.5 Mbps upstream and running five encoders. Mesh tops out around four to six participants before laptops start throttling.
An SFU (Selective Forwarding Unit) collapses that star. The sharer uploads one stream to a WebRTC server, which forwards copies to everyone else. Upload cost stays flat regardless of audience size, and the sharer’s CPU only runs one encoder.
This is what every production conferencing product uses.
An MCU decodes incoming streams, composites them into a single mixed layout, and re-encodes. It’s the most expensive per participant and adds encode-decode latency, but the receiving client only handles one stream, which helps with thin clients and legacy endpoints.
| Topology | Sharer upload | Server cost | Practical ceiling | Added latency |
|---|---|---|---|---|
| P2P mesh | (n-1) × bitrate | STUN/TURN only | 4–6 peers | Lowest |
| SFU | 1 × bitrate | Moderate (routing) | 25–50 with video | Low (~10–50 ms) |
| MCU | 1 × bitrate | High (transcoding) | Hundreds | Higher (100 ms+) |
There’s a fourth option people forget: once your audience passes a few hundred and interactivity stops mattering, switch protocols. A WebRTC vs HLS comparison comes down to whether viewers need to talk back. If they don’t, HTTP delivery scales to millions at a fraction of the cost.
Advantages of WebRTC Screen Sharing
No Plugin, No Extension, No Install
Before the Screen Capture API shipped, Chrome screen sharing required a signed extension and Firefox required a domain allowlist. Both are gone.
WebRTC screen sharing without an extension is now the default path in Chrome, Edge, Firefox, Safari, and Opera on desktop. A plain HTTPS page and a click handler are the whole requirement.
Sub-Second Latency
Screen shares typically reach the far end in 100–300 ms. That’s the difference between a support agent pointing at a button and a support agent describing where the button used to be. Anything built on segmented HTTP delivery adds seconds of video latency before the first frame arrives.
Encrypted by Default
WebRTC mandates DTLS-SRTP. There’s no unencrypted mode to accidentally ship. In a true peer-to-peer session, nothing but the two endpoints can decode the screen contents, which is why regulated industries reach for it over server-side alternatives.
Adaptive Under Bad Networks
Google Congestion Control measures the path continuously and adjusts bitrate and resolution as capacity changes. A share that starts at 1080p on Wi-Fi degrades gracefully on a tethered phone instead of stalling.
Combined with retransmission and forward error correction, it handles the packet loss that kills naive UDP streaming.
One API Across Every Desktop Browser
The same getDisplayMedia() call works in Chromium browsers, Firefox, and Safari on Windows, macOS, and Linux. Browser-specific quirks exist around audio and picker options, but the core capture path is genuinely portable: no per-browser capture backends to maintain.
It Composes With Everything Else in WebRTC
A screen track is just another video track. It rides the same peer connection as camera video and microphone audio, uses the same data channel for cursor coordinates or annotation events, and hits the same getStats() API for monitoring.
Building WebRTC screen sharing with remote control on top means adding a data channel, not a second stack.
Limitations of WebRTC Screen Sharing
Mobile Browsers Can’t Capture at All
No browser on iOS or iPadOS supports getDisplayMedia(), and because Apple requires every iOS browser to use WebKit, “Chrome on iOS” inherits that gap. Android support is patchy.
Screen capture is a desktop web feature. If you need mobile sharing, you need a native app, and React Native WebRTC exposes the platform capture APIs that the browser won’t.
Text Gets Blurry at Default Settings
Encoders tuned for camera video throw away spatial detail to hold frame rate. Applied to a spreadsheet or a terminal, that produces mush. Fixing it means telling the encoder what kind of content it’s handling, which the next section covers.
Audio Capture Is Inconsistent
System audio doesn’t work on macOS or Linux, Safari doesn’t offer it at all, and tab audio is Chromium-only. Any product promising “share your screen with sound” needs per-platform handling and a microphone fallback.
Mesh Topology Collapses Past a Handful of Viewers
The (n-1) upload cost is brutal for screen sharing specifically, because legible screen content needs higher bitrates than a talking-head camera feed. A presenter on a 10 Mbps residential uplink runs out of headroom around five viewers.
Moving to an SFU is the fix, and it means running or renting media servers.
TURN Relay Costs Real Money
Roughly 10–20% of connections can’t establish a direct path and fall back to relay. Every relayed byte is bandwidth you pay for, and screen shares are bandwidth-heavy. Understanding NAT traversal upfront keeps that line item from surprising you after launch.
Those tradeoffs decide the architecture, not the code. With them settled, the build itself is short: here’s the implementation, the quality tuning that makes shared text readable, and the infrastructure you’ll need behind it.
How to Implement WebRTC Screen Sharing
The working version of WebRTC screen sharing in JavaScript is about forty lines. The production version is mostly error handling and lifecycle management.
1. Serve Over HTTPS and Allow display-capture
getDisplayMedia() requires a secure context. Localhost counts during development; everything else needs TLS. If your app runs in an iframe or you’ve enabled Permissions Policy, grant the directive explicitly:
Permissions-Policy: display-capture=(self)
<iframe src="https://app.example.com/share" allow="display-capture"></iframe>
Miss this and the call rejects with NotAllowedError, which looks identical to a user clicking Cancel.
2. Capture the Display Surface
Call getDisplayMedia() from inside a click handler. The options object below picks a sensible default surface, keeps the current tab out of the picker to avoid the hall-of-mirrors effect, and asks for tab audio where it’s available:
const displayMediaOptions = {
video: {
displaySurface: "browser",
frameRate: { ideal: 15, max: 30 },
width: { max: 1920 },
height: { max: 1080 },
},
audio: {
suppressLocalAudioPlayback: true,
echoCancellation: false,
noiseSuppression: false,
},
selfBrowserSurface: "exclude",
surfaceSwitching: "include",
systemAudio: "include",
};
async function startScreenShare() {
try {
const stream =
await navigator.mediaDevices.getDisplayMedia(displayMediaOptions);
return stream;
} catch (err) {
if (err.name === "NotAllowedError") {
console.info("User cancelled the picker or permission was blocked.");
} else {
console.error(`Screen capture failed: ${err.name}`, err);
}
return null;
}
}
Every option here is a hint, not a guarantee.
Constraints apply after the user chooses a surface. They never narrow the picker itself, because that would let a page force someone into sharing their whole desktop. Check track.getSettings() to see what you actually got, and read the full option list in the Screen Capture API docs.
3. Send the Track Over a Peer Connection
If the peer connection already exists and carries camera video, swap the track rather than adding one. replaceTrack() doesn’t trigger renegotiation, so the switch happens in a frame or two instead of a full SDP round trip:
const screenTrack = stream.getVideoTracks()[0];
const sender = pc.getSenders().find((s) => s.track?.kind === "video");
if (sender) {
await sender.replaceTrack(screenTrack);
} else {
pc.addTrack(screenTrack, stream);
}
For a dedicated share alongside camera video, addTrack() is correct. Just expect negotiationneeded to fire, and make sure your signaling handles a mid-call offer.
4. Handle the Browser’s Stop Button
Users stop sharing from the browser’s own indicator, not your UI. That fires ended on the track and nothing else. Miss it and your app shows a “Stop sharing” button for a stream that’s already dead:
screenTrack.addEventListener("ended", async () => {
const cameraTrack = localCameraStream?.getVideoTracks()[0];
if (sender && cameraTrack) {
await sender.replaceTrack(cameraTrack);
}
updateUI({ sharing: false });
});
5. Render the Stream on the Receiving Side
On the far end, ontrack delivers the stream. Setting playsInline and muted on the element avoids autoplay blocks:
pc.ontrack = ({ streams: [remoteStream] }) => {
const video = document.getElementById("remote-screen");
video.srcObject = remoteStream;
video.playsInline = true;
video.muted = true;
};
6. Run Signaling, STUN, and TURN
Nothing above connects on its own. You need three pieces:
- A signaling channel to trade offers, answers, and ICE candidates
- A STUN server to resolve public addresses
- A TURN server for the connections STUN can’t rescue
Public STUN is free. TURN is not, and it’s where screen sharing bandwidth bills accumulate, so budget for relay on 10–20% of sessions.
7. Add Managed Infrastructure for Anything Beyond a Mesh
This is where most projects stall. Peer-to-peer screen sharing between two people is a weekend build. Recording those sessions, reaching a thousand viewers, or delivering to mobile and TV apps is a video infrastructure project measured in months.
The usual fix is to keep WebRTC for the interactive path and hand distribution to a video API. LiveAPI accepts RTMP and SRT ingest, transcodes to adaptive bitrate renditions in seconds, delivers HLS through Akamai, Cloudflare, and Fastly, and records every session automatically.
A screen share that starts as a peer connection can then also be broadcast to a large audience and archived, without you building an origin, a transcoder, or a CDN contract. Teams ship that path in days rather than months.
How to Improve WebRTC Screen Share Quality
A share nobody can read is worse than no share. Four settings do most of the work.
Set contentHint on the Track
contentHint tells the encoder what it’s compressing. Use "detail" or "text" for slides, code, and spreadsheets, and the encoder preserves sharpness at the cost of frame rate. Use "motion" when the shared surface is playing video or animation:
screenTrack.contentHint = "detail";
It’s one line, and it’s the single highest-impact change you can make to a screen share.
Pair It With degradationPreference
When bandwidth or CPU runs short, degradationPreference decides what gives. Match it to the content: maintain-resolution for text and detail, maintain-framerate for motion.
const params = sender.getParameters();
params.degradationPreference = "maintain-resolution";
await sender.setParameters(params);
Setting contentHint to "detail" implicitly sets maintain-resolution when you haven’t specified one, but being explicit avoids surprises across browser versions.
Cap Frame Rate, Not Resolution
Static content doesn’t need 30 fps. Dropping to 5–15 fps frees a large share of the bitrate budget for spatial quality, which is exactly the trade a document share wants. Do it with frameRate: { ideal: 15 } in your constraints, or by setting maxFramerate on the encoding parameters.
Pick the Right Codec
H.264 is the safe interoperability default. VP9 handles screen content noticeably better at the same bitrate because of its screen-content coding tools, and AV1 does better still where both endpoints support it and CPU headroom exists.
You can reorder codec preference per transceiver with setCodecPreferences(), then confirm what actually got negotiated in getStats() instead of trusting the SDP.
Infrastructure You Need for WebRTC Screen Sharing
A Signaling Channel
WebRTC deliberately leaves signaling undefined. Most teams use WebSocket with a small room service that relays SDP and ICE candidates and tracks who’s in a session. Keep it stateless where possible so a reconnect doesn’t lose the room.
STUN and TURN
STUN reveals public addresses. TURN relays when direct paths fail.
Run coturn yourself or buy it. Either way, place relays close to your users, because a TURN server on another continent adds a full round trip to every frame.
A Media Server for Group Sessions
Past four or five participants, you need an SFU. Open source options include mediasoup, Janus, Pion, and LiveKit.
Running one means capacity planning, autoscaling, and monitoring per-stream health, which is a real operations commitment on top of your application code. It’s the same build-or-buy decision teams face when choosing a video conferencing API instead of building the media layer in-house.
Recording and Distribution
Screen shares are worth keeping. Training sessions, incident reviews, and sales demos all have a second life as on-demand content, and live to VOD workflows turn a finished session into a playable asset without a separate export step.
Pair that with multi-CDN HLS delivery and a session that started as a 1:1 peer connection becomes a library anyone in the company can watch later. LiveAPI handles the recording, instant encoding, and global delivery side of that, so your team keeps building the interactive layer instead of an origin and a transcode farm.
Monitoring
getStats() exposes frame rate, bitrate, packet loss, round-trip time, and freeze count per track. Collect it.
Screen sharing failures are usually gradual: a slowly climbing freeze count rather than a hard disconnect. You won’t spot that without telemetry.
Is WebRTC Screen Sharing Right for Your Project?
It’s a good fit if:
- Viewers need to respond in real time, as in support, tutoring, pair programming, and sales demos
- Your users are on desktop browsers
- Sessions involve a handful of participants, or you’re prepared to run an SFU
- End-to-end encryption is a requirement
- You want the interactive path to work without an install
It’s a poor fit if:
- Your primary audience is on phones, where browser capture doesn’t exist
- Thousands of people watch one presenter with no back-channel, since HTTP delivery is far cheaper
- You need broadcast-grade production switching and overlays
- Playback has to reach smart TVs and set-top boxes, which don’t speak WebRTC
Plenty of products need both. WebRTC live streaming carries the interactive core while a streaming API fans the same content out to a large audience, and the handoff between them is where ultra-low latency requirements meet scale requirements.
WebRTC Screen Sharing FAQ
Does WebRTC screen sharing need a browser extension?
No. Chrome dropped the extension requirement when it shipped getDisplayMedia() in version 72, and Firefox dropped its domain allowlist at the same time. Any HTTPS page can request screen capture from a user gesture today.
Can you capture audio along with the screen?
Sometimes. Tab audio works in Chrome, Edge, and Opera. System-wide audio works on Windows and partially on ChromeOS, but not on macOS or Linux, and Safari doesn’t support screen audio at all. Always provide a microphone fallback.
Does WebRTC screen sharing work on mobile?
Not in the browser. No iOS or iPadOS browser implements getDisplayMedia(), and Android support is inconsistent. Native apps can capture the screen through platform APIs and feed those frames into a WebRTC session, but the web path is desktop-only.
How much bandwidth does a screen share use?
A legible 1080p share of mostly static content runs roughly 800 kbps to 2 Mbps. Sharing video or animation at full frame rate pushes past 4 Mbps. In a peer-to-peer mesh, multiply that by the number of viewers to get the sharer’s upload requirement.
Why does shared text look blurry?
The encoder is defaulting to camera-style tuning, which trades spatial detail for frame rate. Set contentHint = "detail" on the video track and degradationPreference = "maintain-resolution" on the sender, then cap frame rate around 15 fps.
How many people can watch a WebRTC screen share?
In a mesh, four to six before the sharer’s uplink and CPU give out. With an SFU, 25 to 50 with video is typical, and a few hundred is achievable with tuning. Beyond that, an HTTP-based protocol is the right tool.
Can you record a WebRTC screen share?
Yes, two ways. MediaRecorder captures the stream client-side into a WebM or MP4 file, which is simple but depends on the sharer’s machine staying online. Server-side recording through a media server or a streaming API is more reliable, and it’s what production systems use.
How do you add screen sharing to a React app?
Wrap getDisplayMedia() in a hook that owns the stream, attaches the ended listener, and cleans up tracks on unmount. Keep the RTCPeerConnection in a ref rather than state so re-renders don’t recreate it. The patterns in a WebRTC React implementation carry straight over to screen capture, since a screen track and a camera track are interchangeable at the peer connection layer.
Final Thoughts
WebRTC screen sharing is two problems wearing one name. Capture is genuinely easy — one method, one picker, a MediaStream, and about forty lines of code.
Delivery is where the engineering lives, and the answer depends almost entirely on how many people are watching and whether they need to talk back.
Get the W3C Screen Capture spec details right on the capture side, set contentHint so text stays readable, respect the screen sharing controls that keep users from oversharing by accident, and pick your topology before you write the signaling layer rather than after.
Ready to add streaming that scales past a peer connection? LiveAPI gives you RTMP and SRT ingest, instant encoding up to 4K, adaptive bitrate HLS delivery across Akamai, Cloudflare, and Fastly, automatic live-to-VOD recording, and pay-as-you-grow pricing — launch in days, not months. Get started with LiveAPI.
