{"id":1361,"date":"2026-10-01T10:28:45","date_gmt":"2026-10-01T03:28:45","guid":{"rendered":"https:\/\/liveapi.com\/blog\/webrtc-sfu\/"},"modified":"2026-10-01T10:29:05","modified_gmt":"2026-10-01T03:29:05","slug":"webrtc-sfu","status":"publish","type":"post","link":"https:\/\/liveapi.com\/blog\/webrtc-sfu\/","title":{"rendered":"WebRTC SFU Explained: How Selective Forwarding Works, Architecture, and When to Use One"},"content":{"rendered":"<span class=\"rt-reading-time\" style=\"display: block;\"><span class=\"rt-label rt-prefix\">Reading Time: <\/span> <span class=\"rt-time\">13<\/span> <span class=\"rt-label rt-postfix\">minutes<\/span><\/span><p>A two-person video call doesn&#8217;t need a server in the middle. A ten-person call does.<\/p>\n<p>Add a third participant to a peer-to-peer call and every browser starts encoding and uploading its video once per person on the call. By six or seven people, most laptops and home connections give up. A WebRTC SFU fixes that by having each participant upload one stream to a server, which forwards it to everyone else.<\/p>\n<p>Google Meet and Discord run on SFUs, and so do the open-source stacks most teams build on: LiveKit, mediasoup, Janus, and Jitsi.<\/p>\n<p>The short version: an SFU keeps each participant&#8217;s upload flat, costs you bandwidth instead of CPU, and stops being the right tool once most of your audience only watches.<\/p>\n<h2>What Is a WebRTC SFU?<\/h2>\n<p><strong>A WebRTC SFU (Selective Forwarding Unit) is a media server that receives audio and video streams from each participant and forwards selected copies of those streams to the other participants, without decoding or re-encoding the media.<\/strong> Each client uploads one stream (or a few simulcast layers) and downloads the streams it wants.<\/p>\n<p>The &#8220;selective&#8221; part matters. The SFU doesn&#8217;t blindly relay everything. It decides, per subscriber, which streams to send and at what quality, based on that subscriber&#8217;s bandwidth, screen layout, and who&#8217;s talking.<\/p>\n<p>The IETF calls this pattern a Selective Forwarding Middlebox in <a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc7667\" target=\"_blank\" rel=\"nofollow\">RFC 7667 on RTP topologies<\/a>, published in November 2015. The industry settled on &#8220;SFU.&#8221;<\/p>\n<table>\n<thead>\n<tr>\n<th>Property<\/th>\n<th>WebRTC SFU<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Full name<\/strong><\/td>\n<td>Selective Forwarding Unit<\/td>\n<\/tr>\n<tr>\n<td><strong>Role<\/strong><\/td>\n<td>Routes RTP media packets between participants<\/td>\n<\/tr>\n<tr>\n<td><strong>Processes media?<\/strong><\/td>\n<td>No decoding or encoding; reads packet headers only<\/td>\n<\/tr>\n<tr>\n<td><strong>Client upload<\/strong><\/td>\n<td>One stream (or 2\u20133 simulcast layers)<\/td>\n<\/tr>\n<tr>\n<td><strong>Client download<\/strong><\/td>\n<td>One stream per visible participant<\/td>\n<\/tr>\n<tr>\n<td><strong>Server cost driver<\/strong><\/td>\n<td>Bandwidth, not CPU<\/td>\n<\/tr>\n<tr>\n<td><strong>Typical latency added<\/strong><\/td>\n<td>Tens of milliseconds<\/td>\n<\/tr>\n<tr>\n<td><strong>Typical session size<\/strong><\/td>\n<td>3 to a few hundred active participants per server<\/td>\n<\/tr>\n<tr>\n<td><strong>Common open-source options<\/strong><\/td>\n<td>LiveKit, mediasoup, Janus, Jitsi Videobridge, Pion-based servers<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>An SFU sits on top of the normal <a href=\"https:\/\/liveapi.com\/blog\/what-is-webrtc\/\" target=\"_blank\">WebRTC<\/a> stack. Clients still use <code>RTCPeerConnection<\/code>, ICE, DTLS, and SRTP. The difference is that the remote peer is a server, not another browser.<\/p>\n<h2>WebRTC SFU vs MCU vs Mesh<\/h2>\n<p>There are three ways to connect more than two WebRTC participants. The choice decides your bandwidth bill, your server CPU bill, and how many people can join.<\/p>\n<h3>Mesh (peer-to-peer)<\/h3>\n<p>In a mesh, every participant opens a direct peer connection to every other participant. No media server is involved, only a <a href=\"https:\/\/liveapi.com\/blog\/webrtc-signaling-server\/\" target=\"_blank\">signaling server<\/a> to exchange SDP and ICE candidates.<\/p>\n<p>It&#8217;s the cheapest to run and the lowest latency. It also falls apart quickly, because each client encodes and uploads its video once per peer.<\/p>\n<h3>MCU (Multipoint Control Unit)<\/h3>\n<p>A multipoint control unit decodes every incoming stream, composites them into a single video (a grid or speaker layout), re-encodes it, and sends each participant one stream.<\/p>\n<p>Clients do little work. The server does a lot: decoding and encoding video for every session is CPU-intensive, adds latency, and locks every viewer into the same layout.<\/p>\n<h3>SFU (Selective Forwarding Unit)<\/h3>\n<p>An SFU splits the difference. Clients upload once like with an MCU, but the server only forwards packets like a router. Each client decodes several streams and lays them out however the app wants.<\/p>\n<h3>Bandwidth math at 10 participants<\/h3>\n<p>Assume each camera stream is 1.5 Mbps and 10 people are on the call.<\/p>\n<table>\n<thead>\n<tr>\n<th>Topology<\/th>\n<th>Upload per client<\/th>\n<th>Download per client<\/th>\n<th>Encodes per client<\/th>\n<th>Server work<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Mesh<\/strong><\/td>\n<td>13.5 Mbps (9 \u00d7 1.5)<\/td>\n<td>13.5 Mbps<\/td>\n<td>9<\/td>\n<td>None<\/td>\n<\/tr>\n<tr>\n<td><strong>SFU<\/strong><\/td>\n<td>1.5 Mbps (about 3.2 Mbps with 3 simulcast layers)<\/td>\n<td>13.5 Mbps, or about 3.7 Mbps with 1 HD + 8 thumbnails<\/td>\n<td>1<\/td>\n<td>Forwards 90 streams<\/td>\n<\/tr>\n<tr>\n<td><strong>MCU<\/strong><\/td>\n<td>1.5 Mbps<\/td>\n<td>~2 Mbps (one composite)<\/td>\n<td>1<\/td>\n<td>Decodes 10, encodes 1\u201310<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Mesh breaks on upload first. Most residential connections can&#8217;t sustain 13.5 Mbps upstream, and a phone can&#8217;t run nine encoders without overheating.<\/p>\n<p>The SFU keeps upload flat no matter how many people join. Download grows with the number of visible participants, but simulcast lets the server send small thumbnails for everyone except the active speaker.<\/p>\n<p>The MCU wins on client bandwidth, but you pay for it in server CPU.<\/p>\n<table>\n<thead>\n<tr>\n<th>Criteria<\/th>\n<th>Mesh<\/th>\n<th>SFU<\/th>\n<th>MCU<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Practical size<\/strong><\/td>\n<td>2\u20134<\/td>\n<td>3 to hundreds (thousands with cascading)<\/td>\n<td>Low tens per server<\/td>\n<\/tr>\n<tr>\n<td><strong>Latency<\/strong><\/td>\n<td>Lowest<\/td>\n<td>Low (adds ~10\u201350 ms)<\/td>\n<td>Higher (decode + encode)<\/td>\n<\/tr>\n<tr>\n<td><strong>Server CPU<\/strong><\/td>\n<td>None<\/td>\n<td>Low<\/td>\n<td>High<\/td>\n<\/tr>\n<tr>\n<td><strong>Server bandwidth<\/strong><\/td>\n<td>None<\/td>\n<td>High<\/td>\n<td>Medium<\/td>\n<\/tr>\n<tr>\n<td><strong>Layout control<\/strong><\/td>\n<td>Client<\/td>\n<td>Client<\/td>\n<td>Server<\/td>\n<\/tr>\n<tr>\n<td><strong>End-to-end encryption<\/strong><\/td>\n<td>Yes<\/td>\n<td>Yes (with Insertable Streams)<\/td>\n<td>No, server must decode<\/td>\n<\/tr>\n<tr>\n<td><strong>Best for<\/strong><\/td>\n<td>1:1 calls<\/td>\n<td>Group calls, webinars, interactive streams<\/td>\n<td>Legacy SIP\/H.323 interop, weak clients<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Most production systems today use an SFU. Some add an MCU-style mixer just for recording or for dialing into phone systems.<\/p>\n<h2>How a WebRTC SFU Works<\/h2>\n<p>A WebRTC SFU looks like a regular peer to each client. Here&#8217;s what happens when someone joins a room.<\/p>\n<ol>\n<li><strong>Signaling:<\/strong> The client connects to the SFU&#8217;s signaling endpoint (usually a WebSocket) and authenticates with a token.<\/li>\n<li><strong>Publish:<\/strong> The client creates an <code>RTCPeerConnection<\/code>, adds its camera and microphone tracks, and sends an SDP offer to the SFU. The SFU answers.<\/li>\n<li><strong>Connectivity:<\/strong> ICE finds a path between the client and the server. Because the SFU has a public IP, this usually succeeds directly, with a <a href=\"https:\/\/liveapi.com\/blog\/turn-server\/\" target=\"_blank\">TURN server<\/a> as fallback for restrictive networks.<\/li>\n<li><strong>Encryption:<\/strong> DTLS runs between the client and the SFU, producing SRTP keys. Media is encrypted on each hop: client to SFU, and SFU to each subscriber.<\/li>\n<li><strong>Forwarding:<\/strong> The SFU receives RTP packets, reads the headers (SSRC, sequence number, timestamp, header extensions), and copies each packet to every subscriber who wants that track. It rewrites sequence numbers and timestamps so each subscriber sees a clean stream.<\/li>\n<li><strong>Subscribe:<\/strong> When a new track is published, the SFU renegotiates with each subscriber and adds the track to their peer connection.<\/li>\n<li><strong>Feedback:<\/strong> Subscribers send RTCP reports (packet loss, NACKs, keyframe requests, bandwidth estimates). The SFU uses them to pick layers, retransmit lost packets, and ask publishers for new keyframes.<\/li>\n<\/ol>\n<p>Steps 5 and 7 are where the &#8220;selective&#8221; happens. The server never decodes video, but it reads enough metadata to switch quality at the right frame boundary.<\/p>\n<h3>Publishers and subscribers<\/h3>\n<p>SFUs model every participant as a publisher, a subscriber, or both. In a group meeting, everyone is both. In a webinar, two or three hosts publish and everyone else only subscribes.<\/p>\n<p>Some SFUs use one peer connection per direction: one for publishing, one for subscribing. Others use a single peer connection with many transceivers.<\/p>\n<p>Both work. Separate connections make renegotiation simpler; a single connection saves one ICE and DTLS handshake.<\/p>\n<h2>Simulcast and SVC: How an SFU Serves Different Networks<\/h2>\n<p>An SFU can&#8217;t transcode, so it needs another way to send a 720p stream to a laptop on fiber and a 180p stream to a phone on 4G. That&#8217;s what simulcast and scalable video coding are for.<\/p>\n<h3>Simulcast<\/h3>\n<p>With simulcast, the publisher encodes the same camera feed two or three times at different resolutions and bitrates, and sends all of them to the SFU. The SFU forwards whichever layer fits each subscriber.<\/p>\n<p>Here&#8217;s how a browser publishes three simulcast layers to an SFU:<\/p>\n<pre><code class=\"language-javascript\">const stream = await navigator.mediaDevices.getUserMedia({ video: true, audio: true });\nconst pc = new RTCPeerConnection({ iceServers: [{ urls: 'stun:stun.l.google.com:19302' }] });\n\nconst [videoTrack] = stream.getVideoTracks();\npc.addTransceiver(videoTrack, {\n  direction: 'sendonly',\n  sendEncodings: [\n    { rid: 'q', scaleResolutionDownBy: 4, maxBitrate: 150_000 },  \/\/ 180p thumbnail\n    { rid: 'h', scaleResolutionDownBy: 2, maxBitrate: 500_000 },  \/\/ 360p\n    { rid: 'f', maxBitrate: 2_500_000 }                           \/\/ 720p full\n  ]\n});\npc.addTrack(stream.getAudioTracks()[0], stream);\n\nconst offer = await pc.createOffer();\nawait pc.setLocalDescription(offer);\n\/\/ Send offer.sdp to the SFU over your signaling channel, then apply its answer\n<\/code><\/pre>\n<p>The <code>rid<\/code> values show up in the <a href=\"https:\/\/liveapi.com\/blog\/webrtc-sdp\/\" target=\"_blank\">SDP<\/a> as <code>a=rid<\/code> and <code>a=simulcast<\/code> lines, defined in RFC 8853. The SFU uses them to tell the layers apart.<\/p>\n<p>Simulcast costs the publisher extra upload (here about 3.15 Mbps instead of 2.5 Mbps) and extra CPU for three encoders. In exchange, the SFU can switch any subscriber between layers instantly.<\/p>\n<h3>Scalable Video Coding (SVC)<\/h3>\n<p>Scalable video coding takes a different approach. The publisher sends one stream built in layers: a base layer plus enhancement layers for higher frame rate (temporal) or higher resolution (spatial). The SFU drops layers to lower the quality for a given subscriber.<\/p>\n<p>SVC is supported in <a href=\"https:\/\/liveapi.com\/blog\/vp9-codec\/\" target=\"_blank\">VP9<\/a> and <a href=\"https:\/\/liveapi.com\/blog\/av1-codec\/\" target=\"_blank\">AV1<\/a>. In Chrome, you request it with <code>scalabilityMode<\/code>:<\/p>\n<pre><code class=\"language-javascript\">pc.addTransceiver(videoTrack, {\n  direction: 'sendonly',\n  sendEncodings: [{ scalabilityMode: 'L3T3' }]  \/\/ 3 spatial, 3 temporal layers\n});\n<\/code><\/pre>\n<p><code>L1T3<\/code> is a common middle ground: one resolution with three temporal layers, so the SFU can drop from 30 fps to 15 or 7.5 fps for weak connections.<\/p>\n<table>\n<thead>\n<tr>\n<th><\/th>\n<th>Simulcast<\/th>\n<th>SVC<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Streams sent by publisher<\/strong><\/td>\n<td>2\u20133 independent streams<\/td>\n<td>1 layered stream<\/td>\n<\/tr>\n<tr>\n<td><strong>Upload overhead<\/strong><\/td>\n<td>~25\u201340% more than a single stream<\/td>\n<td>~10\u201320% more<\/td>\n<\/tr>\n<tr>\n<td><strong>Codecs<\/strong><\/td>\n<td>VP8, H.264, VP9, AV1<\/td>\n<td>VP9, AV1 (H.264 temporal only)<\/td>\n<\/tr>\n<tr>\n<td><strong>Layer switching<\/strong><\/td>\n<td>Needs a keyframe on the target layer<\/td>\n<td>Can switch at layer boundaries<\/td>\n<\/tr>\n<tr>\n<td><strong>Browser support<\/strong><\/td>\n<td>All major browsers<\/td>\n<td>Chromium-based browsers; partial elsewhere<\/td>\n<\/tr>\n<tr>\n<td><strong>SFU complexity<\/strong><\/td>\n<td>Lower<\/td>\n<td>Higher (must parse codec layer metadata)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Most SFUs default to VP8 simulcast because it works everywhere. Teams that control their clients often move to VP9 or AV1 SVC for the bandwidth savings.<\/p>\n<p>Audio is simpler. Everyone sends one <a href=\"https:\/\/liveapi.com\/blog\/opus-codec\/\" target=\"_blank\">Opus<\/a> stream, and the SFU forwards only the loudest few speakers using the audio-level header extension, so a 50-person call doesn&#8217;t push 49 audio streams to every client.<\/p>\n<h2>Bandwidth Estimation and Congestion Control in an SFU<\/h2>\n<p>Picking the right layer depends on knowing each subscriber&#8217;s downstream bandwidth. The SFU has to estimate it continuously.<\/p>\n<p>There are two common methods:<\/p>\n<ul>\n<li><strong>Transport-wide congestion control (TWCC):<\/strong> The receiver reports the arrival time of every packet. The sender (here, the SFU) calculates delay trends and estimates available bandwidth. This is what Chrome and most modern SFUs use.<\/li>\n<li><strong>REMB (Receiver Estimated Maximum Bitrate):<\/strong> The receiver calculates its own estimate and sends one number back. It&#8217;s older and less accurate, but some clients still rely on it.<\/li>\n<\/ul>\n<p>The SFU then runs an allocation step per subscriber: given 3 Mbps of estimated bandwidth, give the active speaker the 720p layer, give eight thumbnails the 180p layer, and pause video for anyone scrolled off screen.<\/p>\n<p>Good SFUs also do:<\/p>\n<ul>\n<li><strong>NACK and retransmission:<\/strong> Resend lost packets from a short buffer instead of asking the publisher.<\/li>\n<li><strong>Keyframe request handling:<\/strong> Collapse many subscribers&#8217; keyframe requests (PLI\/FIR) into one request to the publisher.<\/li>\n<li><strong>Probing:<\/strong> Send padding to test whether a subscriber can handle a higher layer before switching up.<\/li>\n<li><strong>Dynacast \/ pause unused layers:<\/strong> Tell publishers to stop encoding layers nobody is watching, which saves their CPU and upload.<\/li>\n<\/ul>\n<p>This feedback loop is why the same SFU can handle a participant on gigabit fiber and one on a train.<\/p>\n<p>It&#8217;s also the hardest part to build well, and it&#8217;s where SFU implementations differ the most.<\/p>\n<h2>WebRTC SFU Architecture: Single Server, Cascading, and Edge<\/h2>\n<p>One SFU server handles a few hundred active participants before its network interface or CPU fills up. Past that, you need more servers, and how you connect them is the main SFU architecture decision.<\/p>\n<h3>Single SFU per room<\/h3>\n<p>Every participant in a room connects to the same server. A router or load balancer assigns rooms to servers.<\/p>\n<p>This is the simplest design and works for most group meetings. The limits are room size (one server&#8217;s capacity) and geography: a participant in Sydney joining a room hosted in Virginia adds 200+ ms of round-trip time.<\/p>\n<h3>Cascading SFUs<\/h3>\n<p>Participants connect to the nearest SFU, and SFUs relay streams to each other over a backbone. A publisher in London sends to the London server, which forwards one copy to the Singapore server, which fans out to local subscribers.<\/p>\n<p>Cascading cuts latency for global rooms and lets a single session grow past one server. LiveKit calls this a distributed mesh; Jitsi uses Octo relays; mediasoup exposes <code>pipeToRouter<\/code> to connect routers across workers and hosts.<\/p>\n<p>The cost is complexity: you need a control plane to track which tracks live on which node, and every relay hop adds a few milliseconds.<\/p>\n<h3>Edge SFU with anycast<\/h3>\n<p>Some managed providers run SFUs at every point of presence and use anycast so clients connect to the closest one automatically. Cloudflare markets its TURN\/SFU product this way.<\/p>\n<table>\n<thead>\n<tr>\n<th>Architecture<\/th>\n<th>Max session size<\/th>\n<th>Latency for global users<\/th>\n<th>Complexity<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Single SFU<\/strong><\/td>\n<td>Hundreds<\/td>\n<td>Depends on server location<\/td>\n<td>Low<\/td>\n<\/tr>\n<tr>\n<td><strong>Cascading SFUs<\/strong><\/td>\n<td>Thousands to tens of thousands<\/td>\n<td>Low (nearest node)<\/td>\n<td>High<\/td>\n<\/tr>\n<tr>\n<td><strong>Edge\/anycast SFU<\/strong><\/td>\n<td>Thousands+<\/td>\n<td>Lowest<\/td>\n<td>Managed by provider<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>A <a href=\"https:\/\/liveapi.com\/blog\/webrtc-server\/\" target=\"_blank\">WebRTC server<\/a> for a production app usually combines these: a signaling service, an SFU fleet, a TURN fleet, and a control plane that ties rooms to nodes.<\/p>\n<h2>Advantages of a WebRTC SFU<\/h2>\n<h3>Flat upload per participant<\/h3>\n<p>Each client uploads one stream (or one set of simulcast layers) whether the room has 3 people or 300. That&#8217;s what makes group calls possible on home internet and mobile networks.<\/p>\n<h3>Low server CPU<\/h3>\n<p>The SFU never decodes or encodes video. A single server can forward hundreds of streams that would need a rack of GPUs to transcode.<\/p>\n<h3>Low latency<\/h3>\n<p>Forwarding a packet takes microseconds of processing. The SFU adds tens of milliseconds at most, keeping total glass-to-glass delay well under 500 ms. That&#8217;s in the same range as <a href=\"https:\/\/liveapi.com\/blog\/ultra-low-latency-video-streaming\/\" target=\"_blank\">ultra-low latency streaming<\/a>.<\/p>\n<h3>Client-side layout control<\/h3>\n<p>Because each client gets separate tracks, your app decides the layout: grid, speaker view, pinned participants, picture-in-picture. An MCU forces one composite on everyone.<\/p>\n<h3>Per-subscriber quality<\/h3>\n<p>Simulcast and SVC let each subscriber get the quality their connection can handle. One viewer&#8217;s bad Wi-Fi doesn&#8217;t degrade the stream for everyone else.<\/p>\n<h3>End-to-end encryption is possible<\/h3>\n<p>With WebRTC Encoded Transforms (Insertable Streams), clients can encrypt frames before they leave the browser. The SFU still forwards them because it only reads RTP headers.<\/p>\n<h3>Works with data and screen sharing<\/h3>\n<p>The same server can route <a href=\"https:\/\/liveapi.com\/blog\/webrtc-data-channel\/\" target=\"_blank\">data channels<\/a> and <a href=\"https:\/\/liveapi.com\/blog\/webrtc-screen-sharing\/\" target=\"_blank\">screen sharing<\/a> tracks alongside camera feeds, so you don&#8217;t need separate infrastructure for each.<\/p>\n<h2>Limitations of a WebRTC SFU<\/h2>\n<h3>High server bandwidth<\/h3>\n<p>An SFU trades CPU for bandwidth. A 50-person room where everyone watches everyone at 500 kbps pushes about 1.2 Gbps out of the server. Egress is usually the largest line item on an SFU bill.<\/p>\n<h3>Heavy client download and decode load<\/h3>\n<p>Each subscriber decodes multiple streams. Phones struggle past 9\u201316 simultaneous video tiles. You mitigate this with pagination, last-N forwarding (only send the N most recent speakers), and pausing hidden tracks.<\/p>\n<h3>Poor fit for large passive audiences<\/h3>\n<p>An SFU keeps a stateful WebRTC connection, with ICE, DTLS, and congestion control, for every viewer. That&#8217;s fine for 200 people. It&#8217;s expensive and fragile for 50,000. HTTP-based delivery through a CDN is far cheaper per viewer at that scale.<\/p>\n<h3>No built-in composite recording<\/h3>\n<p>Since the SFU never decodes, you can&#8217;t get a single recorded file out of it directly. You need a separate recorder or compositor (often a headless browser or a GStreamer pipeline) that subscribes like any participant.<\/p>\n<h3>Heavy operational overhead<\/h3>\n<p>Running an SFU in production means owning:<\/p>\n<ul>\n<li>TURN servers and regional deployments<\/li>\n<li>Autoscaling based on egress<\/li>\n<li>Per-session monitoring of packet loss and jitter<\/li>\n<li>Fixes for browser updates that change WebRTC behavior<\/li>\n<\/ul>\n<p>It&#8217;s a full-time infrastructure job.<\/p>\n<h2>Open-Source WebRTC SFU Options<\/h2>\n<p>Knowing how an SFU works only gets you so far. The practical questions are which server to use, whether to run it yourself, and what to do when your audience outgrows real-time delivery.<\/p>\n<p>Most teams start with an open-source WebRTC SFU. These are the projects in active development as of October 2026, with figures pulled from GitHub.<\/p>\n<table>\n<thead>\n<tr>\n<th>SFU<\/th>\n<th>Language<\/th>\n<th>License<\/th>\n<th>Latest release<\/th>\n<th>Strengths<\/th>\n<th>Trade-offs<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>LiveKit<\/strong><\/td>\n<td>Go (built on Pion)<\/td>\n<td>Apache-2.0<\/td>\n<td>v1.13.7 (Sept 2026)<\/td>\n<td>Full stack: SDKs for web, iOS, Android, Flutter, React Native; built-in cascading, recording, simulcast\/SVC<\/td>\n<td>Opinionated room model; Redis needed for multi-node<\/td>\n<\/tr>\n<tr>\n<td><strong>mediasoup<\/strong><\/td>\n<td>C++ worker, Node.js\/Rust API<\/td>\n<td>ISC<\/td>\n<td>3.27.1 (Sept 2026)<\/td>\n<td>Fast C++ media path, low-level control over every transport and producer<\/td>\n<td>No signaling or client SDK; you build the app layer<\/td>\n<\/tr>\n<tr>\n<td><strong>Janus<\/strong><\/td>\n<td>C<\/td>\n<td>GPL-3.0<\/td>\n<td>v1.4.2<\/td>\n<td>Plugin system (VideoRoom, Streaming, SIP gateway); mature<\/td>\n<td>GPL licensing; plugin model has a learning curve<\/td>\n<\/tr>\n<tr>\n<td><strong>Jitsi Videobridge<\/strong><\/td>\n<td>Kotlin\/Java<\/td>\n<td>Apache-2.0<\/td>\n<td>Continuous (Jitsi Meet stable builds)<\/td>\n<td>Powers Jitsi Meet; Octo cascading; last-N<\/td>\n<td>Tied closely to the Jitsi stack<\/td>\n<\/tr>\n<tr>\n<td><strong>Pion<\/strong><\/td>\n<td>Go<\/td>\n<td>MIT<\/td>\n<td>v4.2.22 (Sept 2026)<\/td>\n<td>WebRTC library for building your own SFU<\/td>\n<td>A toolkit, not a ready server<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><a href=\"https:\/\/github.com\/livekit\/livekit\" target=\"_blank\" rel=\"nofollow\">LiveKit&#8217;s open-source SFU<\/a> has the largest community (over 21,000 GitHub stars) and the most complete client SDKs. mediasoup is the pick if you want to control every detail and are comfortable writing the signaling layer.<\/p>\n<p>Be careful with older tutorials that recommend ion-sfu. Its last tagged release was in November 2021 and the repository hasn&#8217;t had a push since July 2023.<\/p>\n<h2>Build vs Managed: Running Your Own SFU<\/h2>\n<p>Self-hosting a WebRTC SFU costs nothing in licenses and a lot in engineering time. A production deployment needs:<\/p>\n<ul>\n<li><strong>SFU nodes<\/strong> in each region your users are in, with public IPs and UDP ports open<\/li>\n<li><strong>TURN servers<\/strong> for the 10\u201320% of users behind restrictive firewalls, plus <a href=\"https:\/\/liveapi.com\/blog\/stun-server\/\" target=\"_blank\">STUN<\/a> for address discovery<\/li>\n<li><strong>A signaling, auth, and control plane<\/strong> that issues room tokens, assigns rooms to nodes, and routes cascades<\/li>\n<li><strong>Autoscaling and observability<\/strong> based on bandwidth, plus per-session packet loss, jitter, and freeze metrics<\/li>\n<li><strong>Recording workers<\/strong> if you need archives<\/li>\n<\/ul>\n<p>A managed <a href=\"https:\/\/liveapi.com\/blog\/video-conferencing-api\/\" target=\"_blank\">video conferencing API<\/a> wraps all of that behind an SDK and per-minute pricing. The trade-off is the usual one: control and lower unit cost at scale versus weeks or months of saved engineering time.<\/p>\n<p>If real-time video is your core product, owning the SFU often pays off.<\/p>\n<p>If it&#8217;s one feature among many, a managed service is usually the better bet.<\/p>\n<h2>Scaling Past the SFU: Hybrid WebRTC and HLS Delivery<\/h2>\n<p>An SFU is the right tool when everyone needs to talk. It&#8217;s the wrong tool when a few people talk and thousands watch.<\/p>\n<p>Most large live events use a hybrid design:<\/p>\n<ol>\n<li><strong>Hosts and guests<\/strong> join a WebRTC SFU room for sub-second interaction.<\/li>\n<li><strong>A compositor<\/strong> subscribes to the room and produces one mixed program feed.<\/li>\n<li><strong>That feed<\/strong> is pushed out over RTMP or SRT to a live streaming platform.<\/li>\n<li><strong>The platform<\/strong> transcodes it into adaptive bitrate HLS and delivers it through a CDN to any number of viewers.<\/li>\n<\/ol>\n<p>Viewers get 2\u201310 seconds of latency instead of under one, but delivery cost per viewer drops sharply and playback works on every device, smart TV, and OTT app. Our <a href=\"https:\/\/liveapi.com\/blog\/webrtc-vs-hls\/\" target=\"_blank\">WebRTC vs HLS<\/a> comparison covers the latency and cost trade-offs in detail.<\/p>\n<p>This is where <a href=\"https:\/\/liveapi.com\/live-streaming-api\/\" target=\"_blank\">LiveAPI&#8217;s live streaming API<\/a> fits. You send the program feed from your SFU&#8217;s compositor over RTMP or SRT, and LiveAPI handles encoding, adaptive bitrate HLS output, delivery across Akamai, Cloudflare, and Fastly, and an embeddable player. You don&#8217;t run the SFU fleet for viewers who never need to speak.<\/p>\n<p>LiveAPI also records the stream automatically through <a href=\"https:\/\/liveapi.com\/blog\/live-to-vod\/\" target=\"_blank\">live-to-VOD<\/a>, so the event is available as on-demand video the moment it ends. That solves the SFU&#8217;s recording gap without a separate pipeline.<\/p>\n<p>WHIP is also making this boundary easier to cross. <a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc9725\" target=\"_blank\" rel=\"nofollow\">RFC 9725 (WHIP)<\/a>, published in March 2025, standardizes how a WebRTC sender pushes media to a server with a single HTTP POST, so encoders like OBS can publish straight into an SFU or ingest service.<\/p>\n<h2>Is a WebRTC SFU Right for Your Project?<\/h2>\n<p>You probably need an SFU if:<\/p>\n<ul>\n<li>Your calls regularly have 3 or more participants<\/li>\n<li>Everyone (or most people) needs to speak and be seen<\/li>\n<li>You need total latency under 500 ms<\/li>\n<li>Your app controls the layout on the client<\/li>\n<li>You want end-to-end encryption in group calls<\/li>\n<\/ul>\n<p>You probably need something else if:<\/p>\n<ul>\n<li>Every call is 1:1 (peer-to-peer mesh works, unless you need server-side recording)<\/li>\n<li>You have a few presenters and hundreds or thousands of passive viewers (use HLS through a CDN)<\/li>\n<li>You need to interconnect with SIP or legacy room systems that expect one mixed stream (consider an MCU)<\/li>\n<li>You don&#8217;t want to operate media servers at all (use a managed SDK or a <a href=\"https:\/\/liveapi.com\/blog\/best-live-streaming-apis\/\" target=\"_blank\">live streaming API<\/a>)<\/li>\n<\/ul>\n<p>Many products end up needing both: an SFU for the interactive part and HLS for the audience.<\/p>\n<h2>WebRTC SFU FAQ<\/h2>\n<h3>What does SFU stand for in WebRTC?<\/h3>\n<p>SFU stands for Selective Forwarding Unit. It&#8217;s a media server that receives streams from each participant and forwards them to the others without decoding them. The IETF&#8217;s RFC 7667 calls the same concept a Selective Forwarding Middlebox.<\/p>\n<h3>What is the difference between SFU and MCU?<\/h3>\n<p>An SFU forwards each participant&#8217;s stream individually, so clients receive multiple streams and arrange them locally. An MCU decodes all streams, mixes them into one composite video, and re-encodes it. SFUs use far less server CPU and add less latency; MCUs use less client bandwidth.<\/p>\n<h3>How many participants can a WebRTC SFU handle?<\/h3>\n<p>A single SFU server typically handles a few hundred active participants, depending on bitrate and server bandwidth. Cascading multiple SFUs extends a session to thousands. For tens of thousands of passive viewers, HLS through a CDN is usually cheaper and more reliable.<\/p>\n<h3>Do I need an SFU for a 1:1 WebRTC call?<\/h3>\n<p>No. A 1:1 call works peer-to-peer, with only a signaling server and STUN\/TURN for <a href=\"https:\/\/liveapi.com\/blog\/nat-traversal\/\" target=\"_blank\">NAT traversal<\/a>. You&#8217;d add an SFU for 1:1 calls only if you need server-side recording, transcription, or a consistent media path for monitoring.<\/p>\n<h3>Does an SFU decrypt WebRTC media?<\/h3>\n<p>By default, yes. The SFU terminates DTLS-SRTP on each hop, so it can read the media. To keep the server blind, clients add a second layer of frame encryption with WebRTC Encoded Transforms (Insertable Streams), and the SFU forwards those frames without being able to decode them.<\/p>\n<h3>Is a WebRTC SFU the same as a TURN server?<\/h3>\n<p>No. A TURN server relays packets for a single connection when a direct path isn&#8217;t possible, and it doesn&#8217;t understand media. An SFU is a full WebRTC endpoint that terminates connections, reads RTP headers, picks layers, and fans streams out to many subscribers. Most deployments run both.<\/p>\n<h3>What is the best open-source WebRTC SFU?<\/h3>\n<p>LiveKit is the most popular choice for teams that want a complete stack with client SDKs. mediasoup suits teams that want low-level control in Node.js or Rust. Janus is a mature option with a plugin system, and Jitsi Videobridge is the natural fit if you&#8217;re building on Jitsi Meet.<\/p>\n<h3>Can an SFU stream to platforms like YouTube or an HLS player?<\/h3>\n<p>Not directly, because an SFU only speaks WebRTC. You add a compositor or egress worker that subscribes to the room, mixes the tracks, and pushes RTMP or SRT to a streaming service, which then produces HLS for players and can <a href=\"https:\/\/liveapi.com\/blog\/stream-to-multiple-platforms\/\" target=\"_blank\">restream to multiple platforms<\/a>.<\/p>\n<h2>Picking the Right Path for Your WebRTC SFU<\/h2>\n<p>A WebRTC SFU is the standard way to run group video: one upload per participant, no transcoding on the server, and per-viewer quality through simulcast and SVC. Start with an open-source server like LiveKit or mediasoup, plan for TURN and regional deployments early, and watch your egress bandwidth, since that&#8217;s what you&#8217;ll pay for.<\/p>\n<p>When your audience grows from participants into viewers, don&#8217;t stretch the SFU to cover it. Keep WebRTC for the people on stage and send everyone else an HLS stream through a CDN.<\/p>\n<p>If you want that broadcast side without building encoding and delivery infrastructure, <a href=\"https:\/\/liveapi.com\/\" target=\"_blank\">get started with LiveAPI<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p><span class=\"rt-reading-time\" style=\"display: block;\"><span class=\"rt-label rt-prefix\">Reading Time: <\/span> <span class=\"rt-time\">13<\/span> <span class=\"rt-label rt-postfix\">minutes<\/span><\/span> A two-person video call doesn&#8217;t need a server in the middle. A ten-person call does. Add a third participant to a peer-to-peer call and every browser starts encoding and uploading its video once per person on the call. By six or seven people, most laptops and home connections give up. A WebRTC SFU fixes that [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1362,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_yoast_wpseo_title":"WebRTC SFU: How Selective Forwarding Works, vs MCU and Mesh %%sep%% %%sitename%%","_yoast_wpseo_metadesc":"Learn what a WebRTC SFU is, how selective forwarding, simulcast, and SVC work, how SFU compares to MCU and mesh, and which open-source SFU to pick.","inline_featured_image":false,"footnotes":""},"categories":[31],"tags":[],"class_list":["post-1361","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-webrtc"],"jetpack_featured_media_url":"https:\/\/liveapi.com\/blog\/wp-content\/uploads\/2026\/10\/webrtc-sfu.jpg","yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v15.6.2 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<meta name=\"description\" content=\"Learn what a WebRTC SFU is, how selective forwarding, simulcast, and SVC work, how SFU compares to MCU and mesh, and which open-source SFU to pick.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/liveapi.com\/blog\/webrtc-sfu\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"WebRTC SFU: How Selective Forwarding Works, vs MCU and Mesh - LiveAPI Blog\" \/>\n<meta property=\"og:description\" content=\"Learn what a WebRTC SFU is, how selective forwarding, simulcast, and SVC work, how SFU compares to MCU and mesh, and which open-source SFU to pick.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/liveapi.com\/blog\/webrtc-sfu\/\" \/>\n<meta property=\"og:site_name\" content=\"LiveAPI Blog\" \/>\n<meta property=\"article:published_time\" content=\"2026-10-01T03:28:45+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-10-01T03:29:05+00:00\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\">\n\t<meta name=\"twitter:data1\" content=\"19 minutes\">\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"WebSite\",\"@id\":\"https:\/\/liveapi.com\/blog\/#website\",\"url\":\"https:\/\/liveapi.com\/blog\/\",\"name\":\"LiveAPI Blog\",\"description\":\"Live Video Streaming API Blog\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":\"https:\/\/liveapi.com\/blog\/?s={search_term_string}\",\"query-input\":\"required name=search_term_string\"}],\"inLanguage\":\"en-US\"},{\"@type\":\"ImageObject\",\"@id\":\"https:\/\/liveapi.com\/blog\/webrtc-sfu\/#primaryimage\",\"inLanguage\":\"en-US\",\"url\":\"https:\/\/liveapi.com\/blog\/wp-content\/uploads\/2026\/10\/webrtc-sfu.jpg\",\"width\":1880,\"height\":1253,\"caption\":\"Photo by MART PRODUCTION on Pexels\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/liveapi.com\/blog\/webrtc-sfu\/#webpage\",\"url\":\"https:\/\/liveapi.com\/blog\/webrtc-sfu\/\",\"name\":\"WebRTC SFU: How Selective Forwarding Works, vs MCU and Mesh - LiveAPI Blog\",\"isPartOf\":{\"@id\":\"https:\/\/liveapi.com\/blog\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/liveapi.com\/blog\/webrtc-sfu\/#primaryimage\"},\"datePublished\":\"2026-10-01T03:28:45+00:00\",\"dateModified\":\"2026-10-01T03:29:05+00:00\",\"author\":{\"@id\":\"https:\/\/liveapi.com\/blog\/#\/schema\/person\/98f2ee8b3a0bd93351c0d9e8ce490e4a\"},\"description\":\"Learn what a WebRTC SFU is, how selective forwarding, simulcast, and SVC work, how SFU compares to MCU and mesh, and which open-source SFU to pick.\",\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/liveapi.com\/blog\/webrtc-sfu\/\"]}]},{\"@type\":\"Person\",\"@id\":\"https:\/\/liveapi.com\/blog\/#\/schema\/person\/98f2ee8b3a0bd93351c0d9e8ce490e4a\",\"name\":\"govz\",\"image\":{\"@type\":\"ImageObject\",\"@id\":\"https:\/\/liveapi.com\/blog\/#personlogo\",\"inLanguage\":\"en-US\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/ab5cbe0543c0a44dc944c720159323bd001fc39a8ba5b1f137cd22e7578e84c9?s=96&d=mm&r=g\",\"caption\":\"govz\"},\"sameAs\":[\"https:\/\/liveapi.com\/blog\"]}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","_links":{"self":[{"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/posts\/1361","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/comments?post=1361"}],"version-history":[{"count":1,"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/posts\/1361\/revisions"}],"predecessor-version":[{"id":1363,"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/posts\/1361\/revisions\/1363"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/media\/1362"}],"wp:attachment":[{"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/media?parent=1361"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/categories?post=1361"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/liveapi.com\/blog\/wp-json\/wp\/v2\/tags?post=1361"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}