Conclusion#
This concluding chapter is divided into two parts. Section 8.1 summarizes the contributions made in the previous chapters. Section 8.2 details several opportunities for future research.
Summary of Contributions#
This dissertation proposed machine-oriented compression: a design philosophy, attendant principles, and techniques for compression systems whose signal consumers are machine perception models. Table 1 paired each limitation of conventional codecs with the design principle that addresses it. Chapters 2–7 developed these principles, proceeding from measurement, to codec design, to complete systems.
Chapter 2 established how existing ML systems are impacted by lossy compression, by measuring the machine perceptual quality of compressed signals using task-specific performance metrics for downstream machine perception systems. Across image classification, image segmentation, speech recognition, and music source separation, and across conventional, neural, and generative codecs, three findings emerged: (1) using generative compression, it is feasible to use highly compressed data while incurring a negligible impact on machine perceptual quality; (2) machine perceptual quality correlates strongly with deep similarity metrics; and (3) using lossy compressed datasets (e.g., ImageNet) for pre-training can lead to counter-intuitive scenarios where lossy compression increases machine perceptual quality rather than degrading it.
Chapter 3 introduced WaLLoC, the first machine-oriented codec design, which combines the benefits of conventional transform coding with end-to-end learned compression using neural networks. WaLLoC places a shallow, asymmetric autoencoder and entropy bottleneck between an invertible wavelet packet transform and its inverse, with an encoder that consists almost entirely of linear operations. For RGB images, WaLLoC achieves nearly 6\(\times\) higher compression ratio than the VAE used in Stable Diffusion 3, despite offering a higher degree of dimensionality reduction and similar quality, without perceptual or adversarial losses. Using codecs adopting this design, the chapter proposed a simple but universal strategy for compressed-domain learning, demonstrated on image classification, colorization, document understanding, and music source separation, with improved resolution scaling properties compared to large patch transformers or highly strided CNNs.
Chapter 4 introduced LiVeAction, which increases the performance, accessibility, and flexibility of encoding-efficient asymmetric autoencoders (EE-AAEs) through FFT-like structured operations in the analysis transform, linear attention in the synthesis transform, and a variance-based rate penalty that replaces adversarial and perceptual losses. The resulting modality-agnostic architecture was used to train codecs for stereo and multi-channel spatial audio, RGB and hyperspectral images, video, and 3D medical images.
Chapter 5 introduced FRAPPE, a residual autoencoding framework that uses the full input to predict the residual output via a projection-pursuit encoder. The encoding objective sorts latent channels by importance, providing variable-rate and progressive compression from a single set of encoder weights, while keeping the analysis path an embarrassingly parallel DAG of independent projections with no decoder in the loop. At high compression ratios (\(\sim\)0.1 bpp), FRAPPE-Image provides higher perceptual quality than AVIF with 47\(\times\) faster encoding, making it capable of real-time 1080p, 30 fps CPU-only encoding.
Chapter 6 extended machine-oriented compression to real-time systems, whose latency constraints have long precluded cloud-based processing. DeDelayed divides computation between a remote model operating on delayed video frames and a local model with access to the current frame; the remote model is trained to make predictions on anticipated future frames, and the two models are jointly optimized with an autoencoder that limits the transmission bitrate. On streaming video segmentation with a round-trip delay of 100 ms, DeDelayed improves performance by 6.4 mIoU compared to fully local inference and 9.8 mIoU compared to remote inference.
Chapter 7 addressed compatibility with existing infrastructure. SEAOTTER pairs a sensor-embedded encoder with a one-time transcode performed in the cloud, which reconstructs the signal and re-encodes it as a standard JPEG file whose color transform and quantization matrices are learned end to end. Across global, dense, and vision-language tasks, the transcode increases accuracy relative to the underlying autoencoder. At a compression ratio of 200:1 and compared to AVIF, SEAOTTER provides \(7\times\) faster encoding, \(3.5\times\) faster decoding, and +8% ImageNet top-1 accuracy, while retaining compatibility with JPEG infrastructure.
Taken together, these chapters demonstrate that when compression systems are designed for machine perception from the start, encoder cost, compression efficiency, and compatibility need not be traded against one another.
Future Work#
Beyond MSE loss in encoder training. To support a wide range of applications, the first generation of machine-oriented codecs proposed in Chapters 3, 4, and 5 focused on MSE reconstruction loss. The generative autoencoders that preceded them formulate the reconstruction loss in a way that intentionally removes detail at the encoder so that new detail can be synthesized at the decoder, prioritizing realism over fidelity, and they rely on mechanisms specific to one type of signal, such as a learned perceptual patch similarity or a pre-trained discriminator network for images. Since the full set of downstream applications cannot be anticipated when a codec is designed, the general-purpose MSE loss was chosen based on a theory of exaptation—the use of a representation for a purpose other than the one for which it was optimized—motivated by the empirical correlations observed in Chapter 2. Once a general-purpose codec exists, specialization becomes a matter of fine-tuning rather than building a codec de novo; Chapter 7 took a first step in this direction by training task-aware transcoding pipelines, but for a frozen encoder.
Other researchers have extensively explored split computing (Kang et al., 2017, Matsubara et al., 2022, Matsubara et al., 2023, Teerapittayanon et al., 2016) and feature coding for machines (Choi and Bajić, 2018, Choi and Bajić, 2022, Duan et al., 2020, ISO/IEC, 2025), in which the transmitted representation is an intermediate activation of a particular downstream model. Extending one or more principles of machine-oriented compression, especially asymmetric architectures, to split computing architectures is one direction for future research. Application-specific losses could lead to improved trade-offs between bitrate and downstream accuracy, but additional machinery could be required to maintain the flexibility of a single encoder model. Some promising approaches worthy of exploration include multi-task objectives of the form \(D_{\text{reconstruction}} + \lambda_1 D_{\text{task}} + \lambda_2 R\), or conditioning mechanisms that allow a single encoder network to have its behavior modulated by the current type of task objective.
Unified decoding and score-based modeling. The codecs of Chapters 3–5 are trained with additive uniform noise in place of quantization, following Ballé et al. (Ballé et al., 2017), who show that, with distortion measured as MSE, the resulting relaxed rate-distortion objective is equivalent to the objective of a variational autoencoder (VAE) (Kingma and Welling, 2014): the generative model is a Gaussian centered on the output of the synthesis transform with variance \((2\lambda)^{-1}\), the prior is the factorized entropy model, and the approximate posterior is a uniform density of unit width centered on the output of the analysis transform. The likelihood term of the variational objective then corresponds to the distortion and the prior term to the rate. Diffusion models are latent variable models trained with the same type of variational bound, in which the approximate posterior is fixed to a Markov chain that gradually adds Gaussian noise to the data (Ho et al., 2020); they can be viewed as a hierarchical VAE in the limit of infinite depth (Kingma et al., 2021). Under a particular parameterization, each term of the bound is a denoising score matching objective at one noise level (Ho et al., 2020, Song et al., 2021). An MSE-optimized autoencoder trained with uniform noise and a score-based generative model are therefore two instances of one variational framework, differing in the noise distribution and the number of latent levels.
In LiVeAction, we explored score-based generative enhancement after training an MSE-optimized autoencoder, following a similar procedure adopted by Hoogeboom et al. (Hoogeboom et al., 2023), in which the bitrate is determined entirely by the MSE-optimized autoencoder and a score-based decoder conditioned on its output approximates sampling from the posterior distribution of signals given that output. Designing a decoder and training procedure that could fully unify these two steps from first principles provides an opportunity for greater performance and flexibility when controlling the trade-off between preservation of details and generative decoding of realistic details.
Flexible entropy coding with accurate rate proxies. Neural codecs achieve their rate–distortion advantage with learned nonlinear transforms and learned entropy models (Ballé et al., 2018, Cheng et al., 2020, Minnen et al., 2018, Yang and others, 2023), but on resource-constrained sensors the encoder budget is the binding constraint. End-to-end learned codecs tune entropy coding offline on representative data, but because their DNN-based analysis transforms dominate the cost of encoding, the entropy coder is designed for compression efficiency rather than computational efficiency. EE-AAEs shift the bottleneck to entropy coding, making its cost a first-order concern. A FRAPPE analysis transform costs roughly 10–100 multiply–accumulates per pixel. At that scale, an entropy coder that spends a few dozen operations per latent sample on prediction and context modeling costs as much as the analysis transform it follows.
The codec of Chapter 5 borrowed standard components: each scale’s integer latent tensor was reshaped into a grayscale plane and compressed with JPEG-LS (Weinberger et al., 2000), using \(\log_2\operatorname{Std}(\tilde{z})\) as a rate proxy in the training objective. Both choices were well suited to full-frame image compression; both fail the goals of a general-purpose learned entropy coder. JPEG-LS is a 2-D image codec: applying it to 1-D audio latents requires tiling channel\(\times\)time matrices as two-dimensional planes, mixing unrelated channels under the codec’s vertical prediction, and the strategy does not extend to volumetric signals. Its global context state also means a region’s bits cannot be located or decoded in isolation, precluding per-region detail adaptation. The log-variance proxy is well motivated for the marginal distribution of raw latents: for a generalized Gaussian distribution (GGD) of fixed shape, \(\log_2\operatorname{Std}\) equals the per-sample rate up to an additive constant. But once the coder performs prediction and context modeling, the proxy contains no representation of either, and \(\log_2\operatorname{Std}\) is unbounded below, rewarding variance reduction long after real rate has reached its floor. The coder and rate proxy should therefore be designed together.
An entropy coder for this setting must satisfy four requirements. (1) Implementation cost: it must encode and decode with the precomputed lookup tables, shifts, and compares available on the FPGA and microcontroller encoders for which these codecs are designed; adaptive multialphabet arithmetic coders, whatever their rate advantages, require multiplications and probability-estimate updates at every symbol (Marpe et al., 2003), and ANS-style coders, though they reach symbol-code throughput, carry coder state through every symbol and are decoded in the reverse of encoding order (Duda, 2013). (2) Generality: one mechanism must serve the full range of signals the autoencoders serve, and must code any region of a signal independently of the rest, because the surrounding systems are built on partial decoding, mixed-detail regions, and low-latency streaming. (3) Trainability: the coder must expose a differentiable rate that the autoencoder can optimize against during training, and that rate must reflect the coder actually deployed, prediction and context modeling included. (4) Offline-only adaptation: every parameter is fit before deployment and the decoder loads only fitted scalars—no runtime fitting, no per-context counters, no bitstream-carried model updates.
The latent spaces these encoders produce have properties that simplify that design: bounded integer alphabets, importance-ordered channels, and prediction residuals that empirically follow generalized Gaussian distributions (Sharifi and Leon-Garcia, 1995). One candidate is a GGD model of linear predictive coding residuals, which would retain the coding structures of media standards—precomputed canonical Huffman tables and FIR linear prediction—and replace their design-time assumptions with a small number of parameters learned offline. If the code tables are generated from a closed-form conditional density, the coder’s ideal rate is a differentiable formula of its latents—same prediction, same statistic, same models—enabling direct optimization of the rate–distortion trade-off during encoder training.
Region-adaptive bit allocation. FRAPPE allows progressive reconstructions, but the detail level is chosen globally across the entire image. Modern image and video codecs partition the signal into macroregions and use runtime rate-distortion optimization to efficiently allocate more bits to regions where they favorably contribute to the rate-distortion Lagrangian, and fewer bits where they provide diminishing returns (Bross et al., 2021, Sullivan et al., 2012, Wiegand et al., 2003). Extending the variable-rate capabilities of FRAPPE to a decoder that can accept different detail levels for different regions of the signal provides another avenue to increase compression efficiency and flexibility. Such a decoder depends on an entropy coder that can code each region independently of the rest, which is requirement (2) of the preceding paragraph.
References#
Johannes Ballé, Valero Laparra, and Eero P. Simoncelli. End-to-end optimized image compression. In International Conference on Learning Representations. 2017. URL: https://openreview.net/forum?id=rJxdQ3jeg.
Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. ICML, 2018.
Benjamin Bross, Ye-Kui Wang, Yan Ye, Shan Liu, Jianle Chen, Gary J Sullivan, and Jens-Rainer Ohm. Overview of the versatile video coding (VVC) standard and its applications. IEEE Transactions on Circuits and Systems for Video Technology, 31(10):3736–3764, 2021.
Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. In CVPR. 2020.
Hyomin Choi and Ivan V Bajić. Deep feature compression for collaborative object detection. In 2018 25th IEEE International Conference on Image Processing (ICIP), 3743–3747. IEEE, 2018.
Hyomin Choi and Ivan V Bajić. Scalable image coding for humans and machines. IEEE Transactions on Image Processing, 31:2739–2754, 2022.
Lingyu Duan, Jiaying Liu, Wenhan Yang, Tiejun Huang, and Wen Gao. Video coding for machines: a paradigm of collaborative compression and intelligent analytics. IEEE Transactions on Image Processing, 29:8680–8695, 2020.
Jarek Duda. Asymmetric numeral systems: entropy coding combining speed of huffman coding with compression rate of arithmetic coding. arXiv preprint arXiv:1311.2540, 2013.
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33. 2020. URL: https://arxiv.org/abs/2006.11239.
Emiel Hoogeboom, Eirikur Agustsson, Fabian Mentzer, Luca Versari, George Toderici, and Lucas Theis. High-fidelity image compression with score-based generative models. arXiv preprint arXiv:2305.18231, 2023. URL: https://arxiv.org/abs/2305.18231.
Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang. Neurosurgeon: collaborative intelligence between the cloud and mobile edge. ACM SIGARCH Computer Architecture News, 45(1):615–629, 2017.
Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. In Advances in Neural Information Processing Systems, volume 34. 2021. URL: https://arxiv.org/abs/2107.00630.
Diederik P. Kingma and Max Welling. Auto-encoding variational Bayes. In International Conference on Learning Representations. 2014. URL: https://arxiv.org/abs/1312.6114.
Detlev Marpe, Heiko Schwarz, and Thomas Wiegand. Context-based adaptive binary arithmetic coding in the h. 264/avc video compression standard. IEEE Transactions on circuits and systems for video technology, 13(7):620–636, 2003.
Yoshitomo Matsubara, Marco Levorato, and Francesco Restuccia. Split computing and early exiting for deep learning applications: survey and research challenges. ACM Computing Surveys, 55(5):1–30, 2022.
Yoshitomo Matsubara, Ruihan Yang, Marco Levorato, and Stephan Mandt. Sc2 benchmark: supervised compression for split computing. Transactions on Machine Learning Research, 2023.
David Minnen, Johannes Ballé, and George D Toderici. Joint autoregressive and hierarchical priors for learned image compression. NeurIPS, 2018.
Kamran Sharifi and Alberto Leon-Garcia. Estimation of shape parameter for generalized Gaussian distributions in subband decompositions of video. IEEE Transactions on Circuits and Systems for Video Technology, 5(1):52–56, 1995.
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations. 2021. URL: https://arxiv.org/abs/2011.13456.
Gary J Sullivan, Jens-Rainer Ohm, Woo-Jin Han, and Thomas Wiegand. Overview of the high efficiency video coding (HEVC) standard. IEEE Transactions on Circuits and Systems for Video Technology, 22(12):1649–1668, 2012.
Surat Teerapittayanon, Bradley McDanel, and Hsiang-Tsung Kung. Branchynet: fast inference via early exiting from deep neural networks. In 2016 23rd international conference on pattern recognition (ICPR), 2464–2469. IEEE, 2016.
Marcelo J Weinberger, Gadiel Seroussi, and Guillermo Sapiro. The loco-i lossless image compression algorithm: principles and standardization into jpeg-ls. IEEE Transactions on Image processing, 2000.
Thomas Wiegand, Gary J Sullivan, Gisle Bjontegaard, and Ajay Luthra. Overview of the H.264/AVC video coding standard. IEEE Transactions on Circuits and Systems for Video Technology, 13(7):560–576, 2003.
Y. Yang and others. Computationally-efficient neural image compression with shallow decoders. In ICCV. 2023.
ISO/IEC. ISO/IEC 23888: MPEG Artificial Intelligence (MPEG-AI). 2025. Parts: Part 2: Video coding for machines (VCM), Part 3: Optimization of encoders and receiving systems for machine analysis of coded video content, Part 4: Feature coding for machines (FCM), Part 5: AI-based point cloud coding.