LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

34th ACM International Conference on Multimedia 2026 (ACM MM 2026)

10–14 November 2026

Rio de Janeiro, Brazil

[PDF]

Yiying Wei (AAU, Austria), Xuanhong Chen (Shanghai Jiao Tong University, China), Hadi Amirpour (AAU, Austria) and Christian Timmerer (AAU, Austria)

Abstract: Despite yielding higher visual quality than image-to-image approaches, masked generation paradigms for video face editing fundamentally lacks attribute consistency (e.g., illumination, background). We introduce LumaID, a novel framework that explicitly disentangles identity and expression representations from environmental contexts, enabling high-fidelity, fine-grained video head editing while strictly preserving these crucial attributes. At its core, LumaID employs an Omni-Disentangled Diffusion Transformer (OD-DiT) that leverages 3D proxy representations to thoroughly isolate the source and target facial features, fundamentally preventing identity leakage and illumination degradation. To further overcome the distributional drift caused by proxy estimation noise and the lack of explicit consistency supervision, we propose Consist-GRPO. This post-training reinforcement learning mechanism formulates multi-dimensional reward signals (spanning identity, expression, pose, and lighting) to continuously steer the generative process toward strict spatiotemporal alignment. Extensive evaluations demonstrate that LumaID serves as a highly competitive baseline, exhibiting strong performance over prior approaches in both attributes consistency and overall visual quality.

Selective Multi-Pass Encoding for Cost-Efficient Video Streaming

International Broadcasting Convention (IBC)

[PDF]

Mohammad Ghasempour (AAU, Austria), Hadi Amirpour (AAU, Austria), and Christian Timmerer (AAU, Austria)

Abstract: As video streaming scales across platforms, resolutions, and devices, encoding efficiency has become critical to maintaining quality while controlling computational cost and energy consumption. Multi-pass encoding is widely used in streaming workflows to improve compression efficiency, rate-control accuracy, and quality consistency. However, its computational overhead is applied uniformly across all content, even when additional passes deliver minimal benefit. At scale, this results in unnecessary processing, higher computational cost, and increased energy consumption. This paper presents a content-adaptive strategy, CASE, that predicts whether additional encoding passes would provide meaningful gains using a lightweight mechanism that derives spatial and temporal features from each video segment. We evaluated the approach in two production-oriented scenarios using local and cloud-based video encoders. Results show that the method reduces computational time and encoding cost, with minimal impact on compression efficiency and visual quality. Experimental results show that CASE reduces encoding time by 25.3% on average with only a 2.23% bitrate increase, while its preprocessing and decision overhead is about 1155 times lower than multi-pass encoding time.

LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Mohammad Ghasempour (AAU, Austria), Yiying Wei (AAU, Austria), Hadi Amirpour (AAU, Austria), Christian Timmerer (AAU, Austria)

Abstract: The growing integration of vision and language models is driving a fundamental shift in video understanding and processing. This evolution calls for datasets that jointly capture visual content and its semantic representations at scale. To address this need, we introduce LMM-10K, a large-scale, curated multimodal dataset comprising 10,000 high-fidelity 4K video sequences at 60 fps with rich semantic and perceptual annotations. We developed an automated acquisition pipeline to curate videos from the Pexels repository, using targeted search queries and strict filtering criteria to capture a wide range of real-world scenes. Beyond the video sequences, LMM-10K is enriched with comprehensive multimodal annotations that integrate low-level visual features with high-level semantic information. These include LLM-generated semantic descriptors, no-reference quality metrics, spatial-temporal complexity metrics, and visual diversity attributes. By combining structured annotations with high-quality video data, LMM-10K provides a versatile resource for a wide range of applications, including video enhancement, content-aware compression and streaming, neural video coding, multimodal learning, generative video modeling, and perceptual quality modeling. Dataset URL: Link

Christian Timmerer (AAU/Bitmovin, Austria)

Abstract: The next inflection in streaming isn’t a codec — it’s three pressures (efficiency, low-latency live, and AI-native media) converging on the systems layer, with concrete implications for the MPEG Systems and ITU-T SG21 roadmap drawn from Bitmovin and the ATHENA lab.

ITU-T SG21 and ISO/IEC JTC 1 SC 29 Joint Workshop on “Media Streaming Service – What’s next”

Geneva, July 14, 2026

Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3

 

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Emanuele Artioli (AAU, Austria), Philipp Fößl (AAU, Austria), Shao-Yang Hung (National Tsinghua University, Taiwan), Philipp Fößl (AAU, Austria), Daniele Lorenzi (Bitmovin, Austria), Farzad Tashtarian (AAU, Austria),  Mahdi Dolati (Sharif University of Technology, Iran), Cheng-Hsin Hsu (National Tsinghua University, Taiwan), Christian Timmerer (AAU, Austria)

Abstract: Recent advancements in 3D Gaussian Splatting (3DGS) have enabled photorealistic rendering of complex scenes, yet widespread adoption on mobile and Extended Reality (XR) devices is hindered by substantial computational and bandwidth requirements. While existing solutions often focus on model compression for client-side rendering, they still demand significant GPU power, limiting applicability on resource-constrained hardware. We propose TIGAS (Thin-client Interactive Gaussian Adaptive Streaming), a remote rendering framework offloading rasterization to a backend. To bypass the prohibitive latencies connected to fluctuating network conditions, TIGAS streams view-dependent 2D projections to a lightweight web client over QUIC, minimizing head-of-line (HoL) blocking. A dedicated ABR algorithm adapts rendering quality to fluctuating network conditions, maintaining motion-to-photon latency within strict 6DoF interactive constraints. Furthermore, we discuss the integration of an experimental WebGPU super-resolution pipeline to analyze the trade-offs between perceptual quality enhancements and thin-client processing bottlenecks. We extensively evaluate TIGAS across multi-continental environments using 14 3DGS models and real 6DoF EyeNavGS movement traces. Powered by a backend rendering frames in under 10 milliseconds, TIGAS maintains latency within interactive thresholds while achieving an average SSIM of 0.88, serving both as a robust testbed for 3DGS streaming research and a capable delivery system.

From June 18-19, Dr. Felix Schniz participated in an Austria-wide networking event for scholars working in the field of game studies. Organized by the FG Game Studies, the event brought together researchers from various institutions across the country, including the University of Vienna’s Game Lab, departments of history and sports sciences in Innsbruck, as well as independent researchers.

The event clearly highlighted two main points: the interdisciplinary nature of game studies within larger academic frameworks, and the exemplary role of the University of Klagenfurt in national game research. As the only institution with a dedicated curriculum explicitly focused on games and a strong hybridisation of technical and humanities topics, the work of Dr. Felix Schniz and his colleagues contributing to the Master’s Program in Game Studies and Engineering was recognized as pioneering and set a strong precedent for the future development of game research in Austria.

In addition to exchanging best practices and discussing local and national strategies, the formation of an Austrian research network dedicated to game and play-related studies was also on the agenda. Dr. Felix Schniz expressed his intention to organize a future meeting to bring together interested members.

 

Hadi

On 25.06.2026, Hadi Amirpourazarian defended his habilitation thesis, “The Predictive Video Encoding Using Visual Complexity Analysis”

Congratulations!

Committee members:

Prof. Wolfgang Faber (Chairperson), Prof. Eckehard Steinbach (External Member), Prof. Wilfried Elmenreich, Prof. Barbara Kaltenbacher, Katharina Stengg, Yuliia Lomonosova, Christoph Rauter

Dr Felix Schniz participated in the podcast “Rock my Worlds of English” to promote the Master’s Programme in Game Studies and Engineering.

The full podcast: https://open.spotify.com/episode/4f4wSWb8SD48yFIsTDnG1I?si=Ha53h2m8Txe1LZuiTjJ-EA

 

Hadi

Title: GNS-GAN: A novel GAN model based on gradient noise suppression

Authors: Hongyou Chen, Lingfeng Qu, Baodan Tian, Yutong He, Yong Fan, Hadi Amirpour, Christian Timmerer and Yao Xin

Journal: Applied Soft Computing

Abstract: Generative adversarial networks (GANs) are widely applicable generative models. However, ensuring stability in adversarial learning remains a significant challenge in current GAN training. Gradient noise, among other factors, significantly impacts the stability of adversarial learning in GAN training. To improve the stability of adversarial learning, a gradient noise suppression generative adversarial network model (GNS-GAN) is proposed. This novel GAN addresses gradient noise by establishing stochastic differential equations (SDEs) for gradient noise in both the discriminator and the generator. The factors affecting the stability of adversarial learning are then analyzed using the assumed gradient noise distribution. Subsequently, an adversarial learning method is designed for the discriminator and generator to suppress gradient noise, thereby completing the adversarial training of GNS-GAN. To verify the performance of GNS-GAN, the experimental results are compared and analyzed using CELEBA, BEDROOM, and CIFAR10 datasets. The FID (Fréchet Inception Distance) values are 23.04 for CELEBA, 18.04 for BEDROOM, and 26.59 for CIFAR10. The GNS-GAN model has stable training performances in the tested datasets. These results demonstrate that the novel GAN model enhances the stability of adversarial learning and the quality of the generated images.

On 10 June 2026, Dr Felix Schniz hosted a session on the video game Bloodborne for AAU’s Media Club. Following this semester’s Media Club leitmotif of ‘the fantastic,’ Felix delved into the game’s depiction of arcane architecture, dream spaces, and the sublime in virtual realms. With 15 attendees and even guests from Salzburg on campus who came by just for this specific date, the session was a fantastic conclusion to this semester’s Media Club schedule.