We present different models for reconstructing HDR images with realistic, perceptually accurate highlights from a single photo. We apply physical principles for our formulation to break up the HDR reconstruction into a set of simpler sub-tasks. IntrinsicHDR Differentiating albedo and shading enables shallow neural networks with dedicated supervision. DisentangledHDR Disentangling the diffuse information from directed light allows training on diverse RAW data. RecurrentHDR The constrained reconstruction of a single exposure step opens up the recurrent recovery of bright HDR highlights for everyday photographs.
Thesis Defense Presentation
Dissertation
Intrinsic Single-Image HDR Reconstruction
We formulate dynamic range extension in the intrinsic domain.
The intrinsic image formation model allows us to separate the reconstruction of reflectance from that of illumination.
Both components react differently to clipping: While the dynamic range of the illumination is reduced and local details in the highlights are lost, the surface colors become desaturated.
By applying the intrinsic model, we approach their reconstruction with separate, specialized networks.
We show that our separation of reflectance and illumination enables more precise supervision and improves performance on a wide variety of photographs.
We can compress the highlight distribution when training the shading network without compromising color accuracy in albedo reconstruction.
The new formulation allows us to recover images with high-fidelity and vivid colors.
Highlight Reconstruction through Diffuse and Directed Light
In this chapter, we focus on the different reflective properties of the diverse objects in natural scenes.
Diffuse surfaces reflect incoming light uniformly in every direction and have soft highlights with moderate brightness.
Their appearance differs substantially from the intense light of specular reflections and active light sources directed towards the camera.
We divide the image into the diffuse image and a residual component containing directed light from specular reflections and visible light sources.
Crucially for our goal, the isolated diffuse image already contains large parts of the color information and spans only a moderate dynamic range due to the absence of bright specular extrema.
Accurate diffuse components can thus be retrieved without problem from widely available RAW photos and do not differ significantly from the ones obtained from true HDR images.
As a result, our disentangled modeling of diffuse and non-diffuse highlight reconstruction recovers a realistic dynamic range while preserving details from the diffuse input.
Recurrent Dynamic Range Extension
Our recurrent approach takes on the extension task from a photographic perspective and lastly break up the dynamic range itself.
Instead of recovering the complex HDR image in a single step, we train a neural network to extend the dynamic range of its input by a single exposure value.
This simpler goal amounts to effectively doubling the luminance range, while our network learns to estimate realistic content in a limited range.
Executing the forward pass recurrently on the previous estimations then extends the dynamic range step-by-step until the desired contrast is reached or the clipped pixels are completely reconstructed.
Our recurrent model of dynamic range extension generalizes to images in the wild, while recovering long-tailed physically-accurate highlights even for challenging specular reflections and active light sources.
BibTeX
@PHDTHESIS{dilleHDRphd,
author={Sebastian Dille},
title={Physically-Based Models of Dynamic Range Extension},
year={2026},
school={Simon Fraser University},
}
Sebastian Dille, Keru Fu, S. Mahdi H. Miangoleh, and Yağız Aksoy
SIGGRAPH Asia, 2026
We present an approach to progressively extend the highlights of an image.
Instead of reconstructing the full dynamic range of a complex scene directly,
we learn a simpler task first:
We extend the dynamic range of an input image by a single exposure value.
Once this is mastered, we retrieve the full HDR image for the scene by executing
our network recurrently, progressively increasing the dynamic range of the input.
Our formulation is agnostic to the input dynamic range and targets a bounded output domain.
This enables us to use widely available RAW images for the reconstruction task and adapt
adversarial losses to construct realistic images.
By incorporating Memory Replay for backpropagation, we can train our network recurrently
over multiple inference stages and reduce reconstruction errors.
As a consequence, our system reconstructs challenging long-tailed HDR scenes robustly
and shows powerful recovery of bright light sources and highlights.
@INPROCEEDINGS{dilleRecurrentHDR,
author={Sebastian Dille and Keru Fu and S. Mahdi H. Miangoleh and Ya\u{g}{\i}z Aksoy},
title={Recurrent Dynamic Range Extension},
booktitle={Proc. SIGGRAPH Asia},
year={2026},
}
The low dynamic range (LDR) of common cameras fails to capture the rich contrast in natural scenes, resulting in loss of color and details in saturated pixels.
Reconstructing the high dynamic range (HDR) of luminance present in the scene from single LDR photographs is an important task with many applications in computational photography and realistic display of images.
The HDR reconstruction task aims to infer the lost details using the context present in the scene, requiring neural networks to understand high-level geometric and illumination cues.
This makes it challenging for data-driven algorithms to generate accurate and high-resolution results.
In this work, we introduce a physically-inspired remodeling of the HDR reconstruction problem in the intrinsic domain.
The intrinsic model allows us to train separate networks to extend the dynamic range in the shading domain and to recover lost color details in the albedo domain.
We show that dividing the problem into two simpler sub-tasks improves performance in a wide variety of photographs.
@INPROCEEDINGS{dilleIntrinsicHDR,
author={Sebastian Dille and Chris Careaga and Ya\u{g}{\i}z Aksoy},
title={Intrinsic Single-Image HDR Reconstruction},
booktitle={Proc. ECCV},
year={2024},
}
Sebastian Dille, Ari Blondal, Sylvain Paris, and Yağız Aksoy
ECCV Workshops, 2024
Class-agnostic image segmentation is a crucial component in automating image editing workflows, especially in contexts where object selection traditionally involves interactive tools.
Existing methods in the literature often adhere to top-down formulations, following the paradigm of class-based approaches, where object detection precedes per-object segmentation.
In this work, we present a novel bottom-up formulation for addressing the class-agnostic segmentation problem.
We supervise our network directly on the projective sphere of its feature space, employing losses inspired by metric learning literature as well as losses defined in a novel segmentation-space representation.
The segmentation results are obtained through a straightforward mean-shift clustering of the estimated features.
Our bottom-up formulation exhibits exceptional generalization capability, even when trained on datasets designed for class-based segmentation. We further showcase the effectiveness of our generic approach by addressing the challenging task of cell and nucleus segmentation.
We believe that our bottom-up formulation will offer valuable insights into diverse segmentation challenges in the literature.
@INPROCEEDINGS{dilleBottomup,
author={Sebastian Dille and Ari Blondal and Sylvain Paris and Ya\u{g}{\i}z Aksoy},
title={A Bottom-Up Approach to Class-Agnostic Image Segmentation},
booktitle={Proc. ECCV Workshop},
year={2024},
}
S. Mahdi H. Miangoleh*, Sebastian Dille*, Long Mai, Sylvain Paris, and Yağız Aksoy
CVPR, 2021
Neural networks have shown great abilities in estimating depth from a single image.
However, the inferred depth maps are well below one-megapixel resolution and often lack fine-grained details, which limits their practicality.
Our method builds on our analysis on how the input resolution and the scene structure affects depth estimation performance.
We demonstrate that there is a trade-off between a consistent scene structure and the high-frequency details, and merge low- and high-resolution estimations to take advantage of this duality using a simple depth merging network.
We present a double estimation method that improves the whole-image depth estimation and a patch selection method that adds local details to the final result.
We demonstrate that by merging estimations at different resolutions with changing context, we can generate multi-megapixel depth maps with a high level of detail using a pre-trained model.
@INPROCEEDINGS{Miangoleh2021Boosting,
author={S. Mahdi H. Miangoleh and Sebastian Dille and Long Mai and Sylvain Paris and Ya\u{g}{\i}z Aksoy},
title={Boosting Monocular Depth Estimation Models to High-Resolution via Content-Adaptive Multi-Resolution Merging},
journal={Proc. CVPR},
year={2021},
}
Published CMPT 461 Projects Supervised by Sebastian
Mingrui Zhao*, Duc Anh Nguyen*, Gnanavel Premnath, Sebastian Dille, S. Mahdi H. Miangoleh, and Yağız Aksoy
SIGGRAPH Asia Posters, 2026
We present VETI, a framework that offers physically motivated image editing while remaining fully anchored to the input photograph. VETI poses photo editing as an interactive process within a rendering engine. It provides direct access to the physical properties of the scene by grounding lighting, object materials, and color palettes in an actual 3D representation of the image. Artists can edit real-world photographs with powerful tools typically reserved for synthetic assets, such as palette swatches, BSDF channels, and light rigs. Reflections and illumination changes are rendered realistically, replacing tedious painting of the approximated effects.
@INPROCEEDINGS{zhaoNguyenVETI,
author={Mingrui Zhao and Duc Anh Nguyen and Gnanavel Premnath and Sebastian Dille and S. Mahdi H. Miangoleh and Ya\u{g}{\i}z Aksoy},
title={VETI: 3D-Guided Visual Effect Tuning on Imagess},
booktitle={SIGGRAPH Asia Posters},
year={2026},
}
Tyrus Tracey, Stefan Diaconu, Sebastian Dille, S. Mahdi H. Miangoleh, and Yağız Aksoy
SIGGRAPH Posters, 2025
We propose an interactive pipeline that enables the seamless integration of a 2D logo into a target image, adapting to the surface geometry and lighting conditions of the scene to ensure realistic appearance.
@INPROCEEDINGS{traceyDiaconuCompositing,
author={Tyrus Tracey and Stefan Diaconu and Sebastian Dille and S. Mahdi H. Miangoleh and Ya\u{g}{\i}z Aksoy},
title={Physically-Based Compositing of {2D} Graphics},
booktitle={SIGGRAPH Posters},
year={2025},
}
Samuel Antunes Miranda*, Shahrzad Mirzaei*, Mariam Bebawy*, Sebastian Dille, and Yağız Aksoy
SIGGRAPH Posters, 2024
Near-infrared imagery offers great possibilities for creative image editing. Lying outside the visual spectrum, the NIR information can effectively serve as a fourth color channel to common RGB.
Compared to the latter, it shows interesting and complementary behavior: its intensity strongly varies with the surface materials in the scene and is less affected by atmospheric perturbations.
For these reasons, NIR imaging has been a long-standing topic of interest in research and its integration has been proven successful for applications like false coloring, contrast enhancement, image dehazing, and purification of low-light images.
Recent developments in smartphone technology have simplified the capturing process, making NIR data readily available for broader use outside the research community.
At the same time, existing tools for NIR processing and manipulation are rare and still limited in functionality.
With many solutions lacking specialized features, the editing process is inefficient and cumbersome, making them prone to generate suboptimal results.
To tackle this issue, we introduce a simple and intuitive photo editing tool that combines RGB and NIR properties, offering functions tailored specifically for the RGB+NIR combination, and granting the user the ability to edit and refine images more creatively.
@INPROCEEDINGS{NIREditing,
author={Samuel Antunes Miranda and Shahrzad Mirzaei and Mariam Bebawy and Sebastian Dille and Ya\u{g}{\i}z Aksoy},
title={Interactive RGB+NIR Photo Editing},
booktitle={SIGGRAPH Posters},
year={2024},
}
Brigham Okano, Shao Yu Shen, Sebastian Dille, and Yağız Aksoy
SIGGRAPH Posters, 2022
Art assets for games can be time intensive to produce.
Whether it is a full 3D world, or simpler 2D background, creating good looking assets takes time and skills that are not always readily available.
Time can be saved by using repeating assets, but visible repetition hurts immersion.
Procedural generation techniques can help make repetition less uniform, but do not remove it entirely.
Both approaches leave noticeable levels of repetition in the image, and require significant time and skill investments to produce.
Video game developers in hobby, game jam, or early prototyping situations may not have access to the required time and skill.
We propose a framework to produce layered 2D backgrounds without the need for significant artist time or skill.
In our pipeline, the user provides segmented photographic input, instead of creating traditional art, and receives game-ready assets.
By utilizing photographs as input, we can achieve both a high level of realism for the resulting background texture as well as a shift from manual work away towards computational run-time which frees up developers for other work.
@INPROCEEDINGS{parallaxBG,
author={Brigham Okano and Shao Yu Shen and Sebastian Dille and Ya\u{g}{\i}z Aksoy},
title={Parallax Background Texture Generation},
booktitle={SIGGRAPH Posters},
year={2022},
}