LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

34th ACM International Conference on Multimedia 2026 (ACM MM 2026)

10–14 November 2026

Rio de Janeiro, Brazil

[PDF]

Yiying Wei (AAU, Austria), Xuanhong Chen (Shanghai Jiao Tong University, China), Hadi Amirpour (AAU, Austria) and Christian Timmerer (AAU, Austria)

Abstract: Despite yielding higher visual quality than image-to-image approaches, masked generation paradigms for video face editing fundamentally lacks attribute consistency (e.g., illumination, background). We introduce LumaID, a novel framework that explicitly disentangles identity and expression representations from environmental contexts, enabling high-fidelity, fine-grained video head editing while strictly preserving these crucial attributes. At its core, LumaID employs an Omni-Disentangled Diffusion Transformer (OD-DiT) that leverages 3D proxy representations to thoroughly isolate the source and target facial features, fundamentally preventing identity leakage and illumination degradation. To further overcome the distributional drift caused by proxy estimation noise and the lack of explicit consistency supervision, we propose Consist-GRPO. This post-training reinforcement learning mechanism formulates multi-dimensional reward signals (spanning identity, expression, pose, and lighting) to continuously steer the generative process toward strict spatiotemporal alignment. Extensive evaluations demonstrate that LumaID serves as a highly competitive baseline, exhibiting strong performance over prior approaches in both attributes consistency and overall visual quality.

Selective Multi-Pass Encoding for Cost-Efficient Video Streaming

International Broadcasting Convention (IBC)

[PDF]

Mohammad Ghasempour (AAU, Austria), Hadi Amirpour (AAU, Austria), and Christian Timmerer (AAU, Austria)

Abstract: As video streaming scales across platforms, resolutions, and devices, encoding efficiency has become critical to maintaining quality while controlling computational cost and energy consumption. Multi-pass encoding is widely used in streaming workflows to improve compression efficiency, rate-control accuracy, and quality consistency. However, its computational overhead is applied uniformly across all content, even when additional passes deliver minimal benefit. At scale, this results in unnecessary processing, higher computational cost, and increased energy consumption. This paper presents a content-adaptive strategy, CASE, that predicts whether additional encoding passes would provide meaningful gains using a lightweight mechanism that derives spatial and temporal features from each video segment. We evaluated the approach in two production-oriented scenarios using local and cloud-based video encoders. Results show that the method reduces computational time and encoding cost, with minimal impact on compression efficiency and visual quality. Experimental results show that CASE reduces encoding time by 25.3% on average with only a 2.23% bitrate increase, while its preprocessing and decision overhead is about 1155 times lower than multi-pass encoding time.

LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Mohammad Ghasempour (AAU, Austria), Yiying Wei (AAU, Austria), Hadi Amirpour (AAU, Austria), Christian Timmerer (AAU, Austria)

Abstract: The growing integration of vision and language models is driving a fundamental shift in video understanding and processing. This evolution calls for datasets that jointly capture visual content and its semantic representations at scale. To address this need, we introduce LMM-10K, a large-scale, curated multimodal dataset comprising 10,000 high-fidelity 4K video sequences at 60 fps with rich semantic and perceptual annotations. We developed an automated acquisition pipeline to curate videos from the Pexels repository, using targeted search queries and strict filtering criteria to capture a wide range of real-world scenes. Beyond the video sequences, LMM-10K is enriched with comprehensive multimodal annotations that integrate low-level visual features with high-level semantic information. These include LLM-generated semantic descriptors, no-reference quality metrics, spatial-temporal complexity metrics, and visual diversity attributes. By combining structured annotations with high-quality video data, LMM-10K provides a versatile resource for a wide range of applications, including video enhancement, content-aware compression and streaming, neural video coding, multimodal learning, generative video modeling, and perceptual quality modeling. Dataset URL: Link

Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3

 

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Emanuele Artioli (AAU, Austria), Philipp Fößl (AAU, Austria), Shao-Yang Hung (National Tsinghua University, Taiwan), Philipp Fößl (AAU, Austria), Daniele Lorenzi (Bitmovin, Austria), Farzad Tashtarian (AAU, Austria),  Mahdi Dolati (Sharif University of Technology, Iran), Cheng-Hsin Hsu (National Tsinghua University, Taiwan), Christian Timmerer (AAU, Austria)

Abstract: Recent advancements in 3D Gaussian Splatting (3DGS) have enabled photorealistic rendering of complex scenes, yet widespread adoption on mobile and Extended Reality (XR) devices is hindered by substantial computational and bandwidth requirements. While existing solutions often focus on model compression for client-side rendering, they still demand significant GPU power, limiting applicability on resource-constrained hardware. We propose TIGAS (Thin-client Interactive Gaussian Adaptive Streaming), a remote rendering framework offloading rasterization to a backend. To bypass the prohibitive latencies connected to fluctuating network conditions, TIGAS streams view-dependent 2D projections to a lightweight web client over QUIC, minimizing head-of-line (HoL) blocking. A dedicated ABR algorithm adapts rendering quality to fluctuating network conditions, maintaining motion-to-photon latency within strict 6DoF interactive constraints. Furthermore, we discuss the integration of an experimental WebGPU super-resolution pipeline to analyze the trade-offs between perceptual quality enhancements and thin-client processing bottlenecks. We extensively evaluate TIGAS across multi-continental environments using 14 3DGS models and real 6DoF EyeNavGS movement traces. Powered by a backend rendering frames in under 10 milliseconds, TIGAS maintains latency within interactive thresholds while achieving an average SSIM of 0.88, serving both as a robust testbed for 3DGS streaming research and a capable delivery system.

Title: Can Swarms Be Trusted? Showcasing Swarm Intelligence and Privacy Preservation Through AR 

Conference: SIMULTECH 2026, Porto, Portugal, 18.-20.07.2026

Authors:  Melanie Schranz, M. Gojkovic, Horia Vulcu, Kseniia Harshina, 

Abstract: Swarm intelligence provides a robust approach for decentralized coordination in nowadays systems, yet its algorithmic principles, like local decision-making, role differentiation, and emergent global behavior are often difficult to convey to individuals without prior experience in swarm-based control. This creates practical barriers when deploying swarm-enabled solutions in domains such as shared electric vehicle charging, energy management, or mobility systems, where engineers, operators, and stakeholders must reliably understand how decentralized processes produce system-level outcomes. To address this challenge, we developed an Augmented Reality (AR) game that operationalizes a swarm model inspired by the Artificial Bee Colony algorithm and exposes key algorithmic elements, including information propagation, neighborhood interactions, and collective resource allocation—Swarm AR. The system also illustrates how decentralization can reduce data concentration, which may support privacy advantages under certain assumptions about information flow and system design, without requiring explicit protection mechanisms. A shared electric vehicle charging scenario serves as a use case to demonstrate load balancing and the necessity of distributed coordination. We evaluate the tool through a mixed-method user study using pre/post quantitative measures and qualitative analysis. Results indicate modest improvements in participants’ understanding of swarm coordination logic, decentralized decision processes, and emergent behavior relevant for infrastructure control. These findings suggest that AR-based interactive visualization can serve as an effective technical aid for communicating, validating, and reasoning about the operational characteristics of self-organizing systems, supporting informed engineering design and deployment of decentralized, privacy-aware coordination strategies.

Hadi

Title: An HEVC-based Known-Plaintext Attack for Video Selective Encryption

Authors: Lingfeng Qu, Chen Chen, Jinghan Xu, Yuan Yuan, Ningxiong Mao, Hadi Amirpour

Publication: Springer Nature

Hadi

Title: Asymmetry-Aware No-Reference Video Quality Assessment via Dual-Region Temporal Modeling

Authors: MohammadAli Hamidi, Hadi Amirpour, Christian Timmerer, Luigi Atzori

Abstract: Saliency and semantic-driven asymmetric encoding enable significant bitrate savings while maintaining a comparable viewing experience. This paper presents a No-Reference (NR) Video Quality Assessment (VQA) model for evaluating Asymmetrically Encoded Videos (AEV), addressing challenges such as varying compression levels, scaling artifacts, and asymmetric encoding strategies. The proposed approach combines compression-aware features derived from Quantization Parameters (QPs) with spatio-temporal perceptual descriptors capturing blur, motion, and temporal consistency. A hybrid regression framework based on XGBoost and Ridge regression is employed, where a weighted ensemble improves overall performance. Experimental results conducted on the dataset provided by the QoMEX VQA-AEV Grand Challenge, evaluated under a Leave-One-Source-Out (LOSO) protocol, show that the proposed method outperforms state-of-the-art NR-VQA models in terms of correlation coefficients (Pearson and Spearman) and root mean square error (RMSE).

Hadi

Title: Asymmetry-Aware No-Reference Video Quality Assessment via Dual-Region Temporal Modeling

Authors: Yeganeh Chatri, Hadi Amirpour

Abstract: Modern content-adaptive video encoding increasingly relies on asymmetric compression, where semantically important regions are preserved at higher quality than background areas. This results in spatially and temporally heterogeneous distortion patterns that challenge conventional no-reference video quality assessment (NR-VQA) models, which typically assume spatial homogeneity.

In this work, we propose a lightweight dual-region NR-VQA framework that explicitly models distortion heterogeneity by jointly analyzing global context and a content-focused region using a shared ResNet-18 backbone with temporal mean aggregation. To address limited training data, a two-stage freeze–unfreeze optimization strategy is employed for stable learning.

Experiments on the QoMEX Grand Challenge dataset show that the proposed method achieves an SROCC of 0.881, the highest among the evaluated NR-VQA baselines in our experiments, including NIQE, BRISQUE, DOVER, and Q-Align. Additional evaluations on KoNViD-1k and LIVE-VQC indicate consistent generalization across datasets. These results highlight that explicit modeling of spatial heterogeneity is an effective and practical design principle for NR-VQA under asymmetric compression scenarios.

Hadi

Title: Quality of Multimedia Experience Meets Machine Intelligence

Authors: Wei Zhou, Hadi Amirpour, Tobias Hossfeld

Abstract: Multimedia systems are evolving towards AI-driven, adaptive services, leading to a natural convergence of QoE and machine intelligence. In this context, machine intelligence can empower QoE through learning-based, context-aware, and semantic-driven modelling and optimization. At the same time, QoE can guide machine intelligence by providing a human-centred objective for AI system design and evaluation; see also [11]. Looking beyond human perception, toward agent-centric and hybrid QoE, future multimedia systems increasingly require unified experience objectives that support human-AI co-experience. QoMEX’26 in Cardiff stands as a major milestone highlighting the convergence of Quality of Multimedia Experience with Machine Intelligence. This column reflects on this evolution and outlines the key challenges ahead.

Hadi

Title: DAP-Adapter: Enhancing Few-Shot CLIP with Dynamically Diverse and Context-Aware Prompt Generation

Authors: Zongjian Li, Hongyou Chen, Lingfeng Qu, Yongjie Zhu, Ya Pan, Baodan Tian, Yong Fan, Hadi Amirpour

Abstract: Contrastive language-image pretraining (CLIP) has demonstrated powerful zero-shot and few-shot classification capabilities by training on large-scale image-text pairs. However, in the CLIP training paradigm, data augmentation strategies are applied primarily to the image inputs, whereas the text prompts remain fixed throughout the training process. Existing approaches typically rely on static text templates or use a limited number of learnable soft prompts with categories, which restricts the expressiveness of the model in capturing category semantics. In this paper, we propose a novel approach called the dynamic attribute prompt adapter (DAP-Adapter), which leverages large language models to generate diverse textual descriptions. Our approach introduces attributes as intermediate bridges that link categories to their specific descriptions. During training, a batch-level dynamic language mode sampling mechanism is adopted in combination with learnable soft prompts to dynamically construct rich text prompts. To further enhance its ability to capture semantics, DAP-Adapter also integrates a nontrainable CLIP adapter. To evaluate the model performance, experiments were conducted on ten datasets. The experimental results demonstrate that the proposed DAP-Adapter outperforms the state-of-the-art Tip-Adapter-F method.