The challenge
Video formats were designed for human consumption, not machine perception. As AI enters live production, powering instant replays, tracking, and adaptive on-screen graphics, this mismatch becomes expensive.
Most broadcast systems still rely on mezzanine formats that deliver every pixel, even when models only need a fraction. Each new inference task often triggers another encode or proxy feed multiplying bandwidth, compute, latency and operational complexity.
This “full-frame tax” makes every AI model look like a new viewer, forcing networks to transport entire frames that will mostly be discarded. The result is redundant processing, inflated cloud and egress costs, and slower model response—undermining the ROI of real-time intelligence.
The solution
To make live production AI native, the data format and the network need to work together.
VC-6 (SMPTE ST 2117-1) provides a hierarchical, compute aware structure that lets AI models decode only what they need, such as specific regions or resolutions, without full frame overhead. A single VC-6 stream can serve multiple models, each receiving just the right portion of data for its task.
swXtch.io uniquely provides the AI-native multicast overlay that makes this scalable. It transports VC-6 streams efficiently across ground and cloud networks, so multiple AI models can subscribe to the same feed, each fetching only the pixels relevant to its task. No redundant encodes, no proxy chains, just one data flow serving many models in real time.
Optional integrations
- Where an intra-only codec is not required, MPEG-5 LCEVC can complement existing codecs, improving visual quality while enabling inference at lower resolutions to improve throughput and reduce latency in hybrid workflows.
Together, VC-6 and swXtch.io form a scalable, standards-based foundation for AI-ready live production.
Results and benefits
Compared to ST 2110 uncompressed workflows:
- Bandwidth reduction: from 1.304Gb/s to 118Mb/s (93% reduction)
- Packet rate: 89.7 kpps to 8.1 kpps (91% reduction)
- Processing time: 32ms to 9.6ms (70% improvement)
- AI latency: 1.5s to 317ms (80% faster)
- Fewer encodes: one VC-6 feed supports multiple inference task
- Sustainability: lower compute, energy, and egress costs
Benchmarks on NVIDIA Blackwell GPUs confirmed that VC-6 decoding, natively accelerated on CUDA, is 58% faster than JPEG and nearly 90% faster than JPEG 2000, validating its efficiency for parallel GPU based AI pipelines. The same capability is also available for CPU and OpenCL implementations, ensuring broad deployment flexibility across data center and cloud environments.
These gains translate directly into lower cost per stream, faster model feedback, and greater scalability across live and cloud environments.
Why it matters
This use case shows how compute aware video formats and adaptive multicast networks can make AI a native part of live production. By aligning the codec and transport layers, operators can move pixels once, process them faster, and remove the full frame tax that slows AI adoption today.
It is a practical path toward sustainable, real time AI integration built on open SMPTE standards, interoperable with existing ST 2110 infrastructures, and ready for the next generation of intelligent live workflows.
Key technologies
Key insights
Get early access
The joint VC-6 + swXtch.io AI-native live production solution will be available in Q1. Contact us to register your interest and explore early integration opportunities.
Talk to our team