(ch:conclusion)=
# Conclusion

This concluding chapter is divided into two parts. [Section 8.1](sec:conc_summary) summarizes the contributions made in the previous chapters. [Section 8.2](sec:conc_future) details several opportunities for future research.

(sec:conc_summary)=
## Summary of Contributions

This dissertation proposed *machine-oriented compression*: a design philosophy, attendant principles, and techniques for compression systems whose signal consumers are machine perception models. {numref}`tbl:limitations_principles` paired each limitation of conventional codecs with the design principle that addresses it. Chapters 2--7 developed these principles, proceeding from measurement, to codec design, to complete systems.

Chapter 2 established how existing ML systems are impacted by lossy compression, by measuring the *machine perceptual quality* of compressed signals using task-specific performance metrics for downstream machine perception systems. Across image classification, image segmentation, speech recognition, and music source separation, and across conventional, neural, and generative codecs, three findings emerged: (1) using generative compression, it is feasible to use highly compressed data while incurring a negligible impact on machine perceptual quality; (2) machine perceptual quality correlates strongly with deep similarity metrics; and (3) using lossy compressed datasets (e.g., ImageNet) for pre-training can lead to counter-intuitive scenarios where lossy compression increases machine perceptual quality rather than degrading it.

Chapter 3 introduced WaLLoC, the first machine-oriented codec design, which combines the benefits of conventional transform coding with end-to-end learned compression using neural networks. WaLLoC places a shallow, asymmetric autoencoder and entropy bottleneck between an invertible wavelet packet transform and its inverse, with an encoder that consists almost entirely of linear operations. For RGB images, WaLLoC achieves nearly 6$\times$ higher compression ratio than the VAE used in Stable Diffusion 3, despite offering a higher degree of dimensionality reduction and similar quality, without perceptual or adversarial losses. Using codecs adopting this design, the chapter proposed a simple but universal strategy for compressed-domain learning, demonstrated on image classification, colorization, document understanding, and music source separation, with improved resolution scaling properties compared to large patch transformers or highly strided CNNs.

Chapter 4 introduced LiVeAction, which increases the performance, accessibility, and flexibility of encoding-efficient asymmetric autoencoders (EE-AAEs) through FFT-like structured operations in the analysis transform, linear attention in the synthesis transform, and a variance-based rate penalty that replaces adversarial and perceptual losses. The resulting modality-agnostic architecture was used to train codecs for stereo and multi-channel spatial audio, RGB and hyperspectral images, video, and 3D medical images.

Chapter 5 introduced FRAPPE, a residual autoencoding framework that uses the full input to predict the residual output via a projection-pursuit encoder. The encoding objective sorts latent channels by importance, providing variable-rate and progressive compression from a single set of encoder weights, while keeping the analysis path an embarrassingly parallel DAG of independent projections with no decoder in the loop. At high compression ratios ($\sim$0.1 bpp), FRAPPE-Image provides higher perceptual quality than AVIF with 47$\times$ faster encoding, making it capable of real-time 1080p, 30 fps CPU-only encoding.

Chapter 6 extended machine-oriented compression to real-time systems, whose latency constraints have long precluded cloud-based processing. DeDelayed divides computation between a remote model operating on delayed video frames and a local model with access to the current frame; the remote model is trained to make predictions on anticipated future frames, and the two models are jointly optimized with an autoencoder that limits the transmission bitrate. On streaming video segmentation with a round-trip delay of 100 ms, DeDelayed improves performance by 6.4 mIoU compared to fully local inference and 9.8 mIoU compared to remote inference.

Chapter 7 addressed compatibility with existing infrastructure. SEAOTTER pairs a sensor-embedded encoder with a one-time transcode performed in the cloud, which reconstructs the signal and re-encodes it as a standard JPEG file whose color transform and quantization matrices are learned end to end. Across global, dense, and vision-language tasks, the transcode increases accuracy relative to the underlying autoencoder. At a compression ratio of 200:1 and compared to AVIF, SEAOTTER provides $7\times$ faster encoding, $3.5\times$ faster decoding, and +8% ImageNet top-1 accuracy, while retaining compatibility with JPEG infrastructure.

Taken together, these chapters demonstrate that when compression systems are designed for machine perception from the start, encoder cost, compression efficiency, and compatibility need not be traded against one another.

(sec:conc_future)=
## Future Work

**Beyond MSE loss in encoder training.** To support a wide range of applications, the first generation of machine-oriented codecs proposed in Chapters 3, 4, and 5 focused on MSE reconstruction loss. The generative autoencoders that preceded them formulate the reconstruction loss in a way that intentionally removes detail at the encoder so that new detail can be synthesized at the decoder, prioritizing realism over fidelity, and they rely on mechanisms specific to one type of signal, such as a learned perceptual patch similarity or a pre-trained discriminator network for images. Since the full set of downstream applications cannot be anticipated when a codec is designed, the general-purpose MSE loss was chosen based on a theory of exaptation---the use of a representation for a purpose other than the one for which it was optimized---motivated by the empirical correlations observed in Chapter 2. Once a general-purpose codec exists, specialization becomes a matter of fine-tuning rather than building a codec de novo; Chapter 7 took a first step in this direction by training task-aware transcoding pipelines, but for a frozen encoder.

Other researchers have extensively explored split computing {cite}`teerapittayanon2016branchynet,kang2017neurosurgeon,matsubara2022split,matsubara2023sc2` and feature coding for machines {cite}`choi2018deepfeaturecompression,choi2022scalableimagecoding,duan2020video,mpegAI2025`, in which the transmitted representation is an intermediate activation of a particular downstream model. Extending one or more principles of machine-oriented compression, especially asymmetric architectures, to split computing architectures is one direction for future research. Application-specific losses could lead to improved trade-offs between bitrate and downstream accuracy, but additional machinery could be required to maintain the flexibility of a single encoder model. Some promising approaches worthy of exploration include multi-task objectives of the form $D_{\text{reconstruction}} + \lambda_1 D_{\text{task}} + \lambda_2 R$, or conditioning mechanisms that allow a single encoder network to have its behavior modulated by the current type of task objective.

**Unified decoding and score-based modeling.** The codecs of Chapters 3--5 are trained with additive uniform noise in place of quantization, following Ballé et al. {cite}`balle2017end`, who show that, with distortion measured as MSE, the resulting relaxed rate-distortion objective is equivalent to the objective of a variational autoencoder (VAE) {cite}`kingma2014auto`: the generative model is a Gaussian centered on the output of the synthesis transform with variance $(2\lambda)^{-1}$, the prior is the factorized entropy model, and the approximate posterior is a uniform density of unit width centered on the output of the analysis transform. The likelihood term of the variational objective then corresponds to the distortion and the prior term to the rate. Diffusion models are latent variable models trained with the same type of variational bound, in which the approximate posterior is fixed to a Markov chain that gradually adds Gaussian noise to the data {cite}`ho2020denoising`; they can be viewed as a hierarchical VAE in the limit of infinite depth {cite}`kingma2021variational`. Under a particular parameterization, each term of the bound is a denoising score matching objective at one noise level {cite}`ho2020denoising,song2021score`. An MSE-optimized autoencoder trained with uniform noise and a score-based generative model are therefore two instances of one variational framework, differing in the noise distribution and the number of latent levels.

In LiVeAction, we explored score-based generative enhancement after training an MSE-optimized autoencoder, following a similar procedure adopted by Hoogeboom et al. {cite}`hoogeboom2023high`, in which the bitrate is determined entirely by the MSE-optimized autoencoder and a score-based decoder conditioned on its output approximates sampling from the posterior distribution of signals given that output. Designing a decoder and training procedure that could fully unify these two steps from first principles provides an opportunity for greater performance and flexibility when controlling the trade-off between preservation of details and generative decoding of realistic details.

**Flexible entropy coding with accurate rate proxies.** Neural codecs achieve their rate--distortion advantage with learned nonlinear transforms and learned entropy models {cite}`balle2018variational,minnen2018joint,cheng2020learned,yang2023computationally`, but on resource-constrained sensors the encoder budget is the binding constraint. End-to-end learned codecs tune entropy coding offline on representative data, but because their DNN-based analysis transforms dominate the cost of encoding, the entropy coder is designed for compression efficiency rather than computational efficiency. EE-AAEs shift the bottleneck to entropy coding, making its cost a first-order concern. A FRAPPE analysis transform costs roughly 10--100 multiply--accumulates per pixel. At that scale, an entropy coder that spends a few dozen operations per latent sample on prediction and context modeling costs as much as the analysis transform it follows.

The codec of Chapter 5 borrowed standard components: each scale's integer latent tensor was reshaped into a grayscale plane and compressed with JPEG-LS {cite}`weinberger2000loco`, using $\log_2\operatorname{Std}(\tilde{z})$ as a rate proxy in the training objective. Both choices were well suited to full-frame image compression; both fail the goals of a general-purpose learned entropy coder. JPEG-LS is a 2-D image codec: applying it to 1-D audio latents requires tiling channel$\times$time matrices as two-dimensional planes, mixing unrelated channels under the codec's vertical prediction, and the strategy does not extend to volumetric signals. Its global context state also means a region's bits cannot be located or decoded in isolation, precluding per-region detail adaptation. The log-variance proxy is well motivated for the marginal distribution of raw latents: for a generalized Gaussian distribution (GGD) of fixed shape, $\log_2\operatorname{Std}$ equals the per-sample rate up to an additive constant. But once the coder performs prediction and context modeling, the proxy contains no representation of either, and $\log_2\operatorname{Std}$ is unbounded below, rewarding variance reduction long after real rate has reached its floor. The coder and rate proxy should therefore be designed together.

An entropy coder for this setting must satisfy four requirements. (1) *Implementation cost*: it must encode and decode with the precomputed lookup tables, shifts, and compares available on the FPGA and microcontroller encoders for which these codecs are designed; adaptive multialphabet arithmetic coders, whatever their rate advantages, require multiplications and probability-estimate updates at every symbol {cite}`marpe2003context`, and ANS-style coders, though they reach symbol-code throughput, carry coder state through every symbol and are decoded in the reverse of encoding order {cite}`duda2013asymmetric`. (2) *Generality*: one mechanism must serve the full range of signals the autoencoders serve, and must code any region of a signal independently of the rest, because the surrounding systems are built on partial decoding, mixed-detail regions, and low-latency streaming. (3) *Trainability*: the coder must expose a differentiable rate that the autoencoder can optimize against during training, and that rate must reflect the coder actually deployed, prediction and context modeling included. (4) *Offline-only adaptation*: every parameter is fit before deployment and the decoder loads only fitted scalars---no runtime fitting, no per-context counters, no bitstream-carried model updates.

The latent spaces these encoders produce have properties that simplify that design: bounded integer alphabets, importance-ordered channels, and prediction residuals that empirically follow generalized Gaussian distributions {cite}`sharifi1995estimation`. One candidate is a GGD model of linear predictive coding residuals, which would retain the coding structures of media standards---precomputed canonical Huffman tables and FIR linear prediction---and replace their design-time assumptions with a small number of parameters learned offline. If the code tables are generated from a closed-form conditional density, the coder's ideal rate is a differentiable formula of its latents---same prediction, same statistic, same models---enabling direct optimization of the rate--distortion trade-off during encoder training.

**Region-adaptive bit allocation.** FRAPPE allows progressive reconstructions, but the detail level is chosen globally across the entire image. Modern image and video codecs partition the signal into macroregions and use runtime rate-distortion optimization to efficiently allocate more bits to regions where they favorably contribute to the rate-distortion Lagrangian, and fewer bits where they provide diminishing returns {cite}`wiegand2003overview,sullivan2012overview,bross2021overview`. Extending the variable-rate capabilities of FRAPPE to a decoder that can accept different detail levels for different regions of the signal provides another avenue to increase compression efficiency and flexibility. Such a decoder depends on an entropy coder that can code each region independently of the rest, which is requirement (2) of the preceding paragraph.

## References

```{bibliography}
:filter: docname in docnames
```
