🧠 AI Promises a Revolution in Video Compression

Over the past forty years, the video compression industry has advanced through evolutionary steps: from the MPEG-1 and MPEG-2 standards through the HD broadcasting era with H.264 (AVC), and on to modern 4K formats based on H.265 (HEVC) and the latest VVC/H.266.

Today, however, as artificial intelligence permeates every aspect of television production, the dawn of a new era is clearly emerging on the horizon—the era of pure neural video coding. A landmark event for the market was the acquisition of AI pioneer Deep Render by technology giant InterDigital, confirming the desire of major players to establish a foothold in the nascent AI-native codec space.

In just five years, the history of video compression has caught up with what was created over the previous forty. The potential of neural networks is immense: they are capable of both solving new challenges, such as automated content analysis, and optimizing traditional data compression workflows. Nonetheless, in the short and medium term, experts do not predict a complete displacement of traditional codecs. The most likely scenario is a hybrid model, where AI services are integrated into classical hardware and software architectures.

Hybrid Model and On-the-Fly Upscaling

At the current stage, the value of AI lies in enhancing the Quality of Experience (QoE) and optimizing operational costs. Utilizing neural networks at the pre-processing stage allows for more precise frame rate selection and locally improves image sharpness. At the post-processing stage, it enables deep, real-time video quality assessment that outperforms traditional metrics like VMAF. AI helps operators find the perfect balance between visual quality for the end user and signal delivery costs.

Another promising avenue is AI upscaling (Super Resolution). A prime example is the technological partnership between Beamr and Nvidia, who demonstrated a pipeline capable of scaling HD content to 4K resolution in real time. This approach leverages the core strengths of a classical codec—spatial-temporal redundancy removal and entropy coding—while delivering an entirely different level of picture quality to consumer devices. This solution bridges the gap caused by the lack of local 4K infrastructure, enabling broadcasters to deliver premium content over standard communication channels.

Barriers to Pure Neural Coding

Despite the obvious benefits, the transition to fully neural coding faces serious technological and ethical challenges:

  • The issue of authenticity and “hallucinations.” Unlike mathematically precise traditional algorithms, generative AI can invent pixels that were not present in the original footage. In live news and sports broadcasting, where the battle for content verification is being fought (within the framework of C2PA standards), any AI-induced distortion could undermine audience trust.
  • Interoperability and standardization. A classical codec can be decoded at any point along the broadcast technology chain. In contrast, AI solutions require end-to-end compatibility. A significant first step in this direction was the approval of the JPEG AI standard for still images, but adapting these approaches to full-motion video while maintaining high frame rates is orders of magnitude more complex.
  • Computational power. Neural video coding is highly resource-intensive for GPUs, particularly during fast-paced broadcasts like sports matches. The cost of deploying such GPU-heavy cloud workflows remains prohibitively high for most media companies.
  • Consumer hardware. The tightest bottleneck is the decoder ecosystem. Any changes requiring consumer set-top boxes or smart TVs to perform complex AI computations demand a rock-solid economic justification.

The Future of the Industry

In traditional broadcasting and cinema, where preserving every detail and the creator’s original artistic intent is critical, classical codecs and hybrid models will continue to dominate for many years to come.

However, in fully automated sectors—such as autonomous vehicles, where massive volumes of video data from cameras are transmitted for analysis by other AI systems—pure neural compression could become the standard in the foreseeable future. The current pipeline of video encoders is sequential, which imposes limitations on parallel computing. AI has the potential to shatter this glass ceiling. The processes will not become simpler, but they will become exponentially faster.

Source: TVB Europe