Research
Highlights
ICPR 20206Best Paper Award for our work on document parsing...
BMVC 262 paper accepted
ECCV 2026full paper + 4 workshop paper accepted
TPAMI 2026full paper accepted
Neu Open Source Lib NCATorch
TMLR 263 journal paper
ICML 2026presenting 2 paper: a full paper and a TMLR J2C
ICLR 262 full paper accepted
TMLR 26with J2C Certification (top 10%)
NeurIPs 25oral presentation and 2 workshop paper
Resources
Videoson YouTubeFind videos from our talks and online paper presentations in our YouTube Channel
Source CodeOn GitHubWe typically publish the code with our papers on our GitHub page.
Publications
Selected list of recent papers. A full list of all publications can be found at our Google Scholar page.
2026
Robustness should extend across classes, rather than concentrating on the easiest ones. RL-FAT combines reinforcement-learning-inspired feedback with a fairness-focused loss to reduce class-wise disparities during adversarial training.
Table extraction should preserve meaning, not just matching strings. This benchmark combines controlled PDFs with an LLM-based semantic assessment that aligns more closely with human judgments than conventional structural metrics.
Ordinary smartphone photos are a demanding test for deepfake detectors. LAION-Mobile assembles roughly one million images with camera metadata and reveals serious generalization and threshold-calibration problems in existing detectors.
Visual concepts can make biological image classification easier to interpret. This work investigates unsupervised concept bottlenecks as a way to connect bioimage predictions with intermediate visual evidence.
Rapidly changing product catalogs challenge static prediction models. This work studies how fine-tuning and retrieval-augmented generation can be combined to extract structured information from multimodal retail data.
Explanations of abstract classifications need a clear theoretical reference. This position argues for evaluating model explanations against validated, context-specific constructs to make model bias measurable.
A tracker that keeps running may still accumulate serious drift. This study separates tracking failure from trajectory degradation and tests when synthetic corruptions reproduce conclusions drawn from real adverse conditions.
Sharp details should survive compression into a generative model's latent space. DeBaT learns low- and high-frequency representations separately, improving fine-detail reconstruction while retaining coherent global structure.
An image-quality metric should notice when a small prompt change changes the intended meaning. CROC generates contrastive checks at scale, adds a human-supervised benchmark, and uses the resulting data to train a more capable evaluation metric.
Event cameras offer a fast, sparse view of moving particles. STELLA unifies alternative detection and tracking pipelines with synthetic and experimental benchmarks for studying detailed motion in fluid flows.
What useful visual information survives directly in low-bit RAW sensor data? RAWDet-7 benchmarks object detection and description across cameras, environments, and quantization levels, bringing sensor constraints into model evaluation.
Lens blur is not always a harmless image degradation. Adverse Lens Corruption optimizes physically parameterized aberrations to find difficult optical conditions and test a model's sensitivity to lens tolerances.
Camera optics introduce blur that simple synthetic kernels do not capture. OpticsBench and LensCorruptions probe realistic aberrations across classification and detection models, supporting more representative robustness evaluation.
A formula can be written differently while retaining exactly the same meaning. This benchmark uses controlled PDF generation and semantic evaluation to compare formula extraction, with human judgments validating the assessment approach.
Hidden instructions in a paper can distort an automated review. Our experiments examine both prompt-injection susceptibility and acceptance bias in LLM-generated reviews, exposing weaknesses in using these systems for scientific evaluation.
Does Fréchet Inception Distance preserve a trustworthy ordering outside familiar image distributions? This work examines rank consistency for synthetic out-of-distribution samples, focusing attention on the reliability of generator evaluation.
Finding the right spare part requires distinguishing nearly identical objects across viewpoints and backgrounds. A new industrial dataset and lightweight adaptation framework test how foundation-model features can improve this demanding retrieval task.
Simple local update rules can produce complex learned behavior. This review organizes neural cellular automata into a unified framework and introduces NCAtorch as a modular reference implementation for reproducible experimentation.
Treat image features as a table and adapt a detector through examples rather than retraining. This approach combines frozen visual features with TabPFN, showing promise when only a few labeled images from a new generator are available.
Modern phone photographs already contain substantial algorithmic processing. This position paper asks what a "real" image means for deepfake detection and calls for definitions and benchmarks that reflect contemporary photography.
RobustSpring tests motion and depth estimation under corruptions that remain consistent across time, viewpoints, and scene depth. Its benchmark complements accuracy scores with robustness measurements, exposing weaknesses that clean images can conceal.
Generative models can show the world through a narrow set of geographic stereotypes. GeoDiv measures socioeconomic portrayal and visual diversity separately, making these biases more interpretable across countries and image generators.
Extracting structured product information demands more than reading text. mSOP-765k supplies over 765,000 annotated advertisement images and evaluation tools for comparing multimodal models, including retrieval-augmented approaches.
2025
Do safety-training gains survive edits to a model's internal activations? This checkpoint-level study tests refusal behavior before and after abliteration and examines how evaluation judges affect the conclusions.
Explore changes to an urban sound map in a fraction of a second. Conditioned normalizing flows learn to predict sound propagation from city layouts, enabling rapid comparison of source and geometry changes.
Label smoothing can unintentionally reinforce mistakes and collapse feature diversity. MaxSup targets the largest prediction logit instead, preserving richer representations while reducing overconfidence more consistently.
AIM encourages models to rely on meaningful object features through self-supervised masking. It improves inherent interpretability alongside classification performance without requiring extra region annotations.
Turn a caption into editable graphics code without requiring aligned caption-program pairs for training. TikZero uses image representations as a bridge between text understanding and graphics-program synthesis, enabling precise, reusable figures.
A different input representation can help vision models cope with darkness. Inspired by retinal processing, fixed color and contrast transformations emphasize structural cues and improve segmentation under difficult lighting.
Fine textures can disappear when image tokenizers favor low-frequency structure. This study diagnoses that imbalance and explores separate optimization of frequency bands to preserve sharper details in reconstructed images.
Specialist imaging tasks often have too few labels for conventional training. We adapt vision-language models to terahertz imagery through modality-aware prompts and in-context examples, exploring classification and interpretation without fine-tuning.
FlowBench brings systematic robustness testing to optical-flow estimation. Its shared evaluation tools compare models under attacks and distribution shifts, helping researchers assess reliability beyond clean benchmark accuracy.
Choosing a text-sampling method means balancing variety against unreliable continuations. This framework evaluates that trade-off at individual decoding steps and offers practical guidance for selecting truncation methods and parameters.
DCBM builds interpretable classifiers from dataset-specific visual concepts. Foundation-model region extraction makes efficient use of limited examples, while localized concepts help explain predictions in fine-grained and unfamiliar domains.
Fast predictions are not enough if a learned simulator gets the physics wrong. PhysicsGen benchmarks generative models on three image-based simulation tasks, making both their speed potential and physical limitations visible.
Add new retail products without retraining the classifier. Our visual RAG pipeline retrieves a few relevant examples to guide a vision-language model in predicting product identity, price, and promotion details.
Generate detailed 3D shapes through compact, multiscale wavelet representations. 3D-WAG predicts progressively finer token maps, reducing the long sequences that make conventional autoregressive 3D generation expensive.
Interpretable classifiers should not need enormous concept collections. Data-efficient visual concept bottlenecks derive concepts from image regions, supporting fine-grained recognition with concepts that can be localized in new images.
Do artificial corruptions tell us how models will behave in real adverse conditions? This large segmentation study finds useful aggregate correlations while showing why individual corruption types still need careful interpretation.
Move an object away from the center or make it smaller, and a classifier may lean more heavily on its background. Hard-Spurious-ImageNet exposes these interactions and tests whether existing bias-mitigation methods cope with them.
Accurate stereo matching is only part of dependable depth perception. DispBench systematically evaluates disparity models under adversarial attacks and image corruptions, making reliability and generalization easier to compare.
Train on simulated seismic data, then remove unwanted multiples in field recordings. This study compares training objectives and shows the value of predicting multiples before subtracting them, including under noisy conditions.
Can instructions change how a model sees an object? We study texture and shape preferences in vision-language models, revealing both the influence of multimodal training and the limits of steering perception with language.
Longer generated videos need meaningful change over time. VSTAR combines a sequence of text prompts with temporal-attention regularization to guide pretrained video models toward more dynamic, evolving scenes.
Average robustness can hide large differences between classes. FAIR-TAT uses targeted adversarial training to improve the balance of class-wise performance and explore fairer trade-offs on clean and perturbed inputs.
Discover recurring visual themes without fixing the number of clusters beforehand. A minimum-cost multicut approach groups images for visual framing analysis and shows how embedding choices expose different levels of thematic detail.
2024
Can an untrained network reveal how robust it will become? This evaluation finds that robustness prediction is harder than clean-accuracy prediction and benefits from combining several zero-cost architecture proxies.
Which layers a network actually needs depends on how it was trained. By resetting layers across differently trained ImageNet models, we reveal substantial changes in where decision-critical information resides.
Explanations need reliable evaluation, too. We replace disruptive pixel deletion with adversarial perturbations to assess attribution maps more consistently and reduce the distribution shifts that can distort their rankings.
A single preprocessing layer can recover information that remains useful under image corruptions. With fewer than 2,000 additional weights, this approach learns a simple linear transformation that improves classification stability at low cost.
Build geometric knowledge directly into an image segmentation. Carefully parameterized implicit representations can enforce properties such as convexity, symmetry, and connectedness, helping resolve difficult or occluded boundaries.
Attention should suppress irrelevant entries without losing several useful alternatives. MultiMax addresses this balance with an adaptive normalization function that preserves multiple modes while encouraging sparsity.
CosPGD provides a common adversarial test for pixel-wise prediction tasks. Its smooth, prediction-alignment weighting produces effective attacks across both classification and regression outputs, from segmentation to optical flow.
How large would convolution filters grow if their size were no longer expensive? Neural Implicit Frequency Filters make this question testable and reveal that many learned filters remain spatially compact even when much larger ones are available.
Language can change which visual cues a multimodal model follows. We measure texture and shape preferences across vision-language models and explore how prompting steers their decisions.
Sometimes a difficult label is ambiguous rather than simply wrong. By studying pedestrian annotations, we show how identifying such cases can improve training efficiency and detection performance while preserving dataset representativeness.
Do familiar visual biases explain why some models generalize better? A controlled study of ImageNet models finds that shape and spectral biases alone cannot reliably predict performance across diverse distribution shifts.
Style synthesis can prepare segmentation models for places and conditions they have never seen. This extended framework mixes content with styles from both training images and external exemplars, and explores stylized validation data for model selection.
Recognizing the same product is different from finding products that serve the same purpose. Retail-786k introduces large-scale visual entity matching with real advertisement images, challenging models to transfer product-equivalence concepts to unseen examples.
Urban sound propagation turns a complex physical process into a concrete test for generative models. This benchmark pairs city layouts with simulated sound maps, revealing where fast learned predictions capture physics and where they fall short.
Generate images that follow a scene layout and remain editable through text. ALDM adds adversarial supervision to diffusion training, improving layout fidelity and making generated scenes useful for segmentation data augmentation.
Top-GAP encourages a classifier to focus on compact, informative image regions. The resulting representations reduce background dependence while improving interpretability, localization, and robustness.
Can one vision-language model replace a carefully engineered retail extraction pipeline? Our case study finds strong performance on some attributes but substantial gaps on fine-grained product identity and discounts, identifying where production challenges remain.
Is an image detector learning synthetic content, or just JPEG compression? We uncover compression and image-size shortcuts in generation-detection datasets and show how removing them changes robustness and cross-generator evaluation.
Upsampling must recover fine detail while keeping predictions stable. We investigate spectral artifacts and show why access to a larger spatial context matters for robust, high-quality pixel-wise outputs.
2023
Encourage robustness directly through the filters a CNN learns. Frequency regularization promotes lower-frequency representations and improves resilience to attacks and distribution shifts without requiring adversarial examples during training.
Seismic multiple removal can be simplified with a U-Net trained entirely on synthetic data. Alongside field-data experiments, this study examines hyperparameters and uncertainty to make the model's behavior easier to understand and use.
Strong image-restoration accuracy can conceal severe adversarial vulnerability. We examine restoration transformers and related architectures, then investigate training and design changes that improve their resistance to attacks.
The border of an image can expose a hidden architectural weakness. We analyze how convolutional padding shapes adversarial perturbations and how alternative padding choices affect robustness.
How can synthetic imagery help assess a segmenter's behavior beyond its training domain? This work investigates synthetic data as a tool for evaluating domain generalization in semantic segmentation.
Change an image's style without changing its semantic layout. Intra-source style augmentation uses a masked-noise StyleGAN encoder to diversify training scenes and improve segmentation under unfamiliar weather and appearance conditions.
What makes one neural architecture more robust than another? This workshop study introduces robustness evaluations across NAS-Bench-201 and demonstrates their use for architecture search, robustness prediction, and analysis of design choices.
Nearly identical products are hard to distinguish from pictures alone. Our leaflet dataset and multimodal classifier show how combining product imagery with extracted text improves fine-grained retail recognition.
Explore the edge of a generator's training distribution by optimizing its discrete latent space. Using smiling faces as an illustrative test case, our approach combines tree-based optimization with weighted retraining to produce samples with stronger target attributes.
Adversarial training can shift image recognition toward the shapes that humans rely on. We examine this effect across architectures and attack settings, using frequency analysis to investigate why more human-like visual preferences emerge.
Can a generative model become a versatile seismic-processing tool? We investigate diffusion models for multiple removal, denoising, and interpolation, comparing their behavior on synthetic and field data with established methods.
Architecture design has a measurable impact on robustness, even at similar parameter counts. This dataset evaluates NAS-Bench-201 architectures under attacks and corruptions, enabling repeatable searches for networks that are both accurate and resilient.
2022
Better robustness does not have to come at the expense of generated-image quality. We regularize deterministic autoencoders using perturbed examples and latent-distribution comparisons to improve both representation stability and synthesis fidelity.
Robustness training can also improve how cautiously a model makes predictions. This study examines the lower overconfidence of adversarially trained networks and the influence of activation functions and pooling on confidence.
Real lenses produce more complex blur than standard corruption benchmarks assume. We evaluate direction- and wavelength-dependent optical effects, exposing classification changes that simpler blur tests can miss.
Do medical images require fundamentally different convolution filters? A closer look shows that apparent outliers largely reflect architectural processing choices, supporting the value of diverse pretraining data across image domains.
Climate models need detailed aerosol physics without prohibitive simulation costs. Our neural emulator accelerates aerosol microphysics while incorporating physical constraints to improve mass conservation and positivity.
Instead of repeatedly searching unpromising architectures, learn where good candidates are likely to be. AG-Net combines a generator with a performance predictor to guide efficient search, including joint optimization of accuracy and hardware latency.
A small change to downsampling can make adversarial training more stable. FrequencyLowCut pooling removes aliasing and helps prevent catastrophic overfitting during fast, single-step adversarial training.
CNN Filter DB makes more than a billion learned filters available for studying what neural networks learn. Its analysis reveals both shared filter statistics across tasks and degenerate filters that can undermine robustness and transfer learning.
Robustness leaves a signature in learned convolution filters. By comparing adversarially trained networks with standard models, we uncover changes in filter diversity, sparsity, and early-layer filtering that help explain their different behavior.
Aliasing reveals when adversarial training starts to lose its robustness. This extended study connects downsampling artifacts to robust overfitting and proposes an aliasing-based early-stopping criterion.
Grouping points into lines, motions, or geometric transformations requires relationships beyond pairs. Our higher-order multicut formulation captures these relationships and provides an efficient local-search solver without fixing the number of groups in advance.
Evaluating architecture-search methods need not require training every candidate network. Learned surrogate benchmarks make large search spaces accessible at a fraction of the cost, supporting more realistic and reproducible NAS comparisons.
Adversarial training changes more than resistance to attacks. Our experiments show that robust models can also be less overconfident on clean inputs, with architecture choices influencing their prediction confidence.
A high robustness score is only useful when the benchmark reflects the intended threat. We examine the detectability and resolution dependence of AutoAttack perturbations, questioning how far common benchmark rankings transfer to practical settings.
Downsampling artifacts offer a revealing clue to adversarial vulnerability. This study shows that robust CNNs learn to downsample more accurately and exhibit less aliasing than their standard counterparts.
2021
Convincing images should have convincing frequency statistics, too. A lightweight spectral discriminator helps GANs match real-image spectra and reduces the frequency artifacts that can expose generated content.
Tell a generator how many objects of each class to include. Our count-conditioned GAN combines image synthesis with object counting, enabling explicit control over scene composition even against complex backgrounds.
What changes inside a CNN when its training data or task changes? This early study compares more than half a billion learned convolution filters, opening a new window onto transfer learning and distribution shifts through model weights.
Shape your Space gives deterministic autoencoders an expressive, multimodal latent distribution during training. This makes sampling straightforward while avoiding a separate density-fitting step after training.
Searching for better networks starts with a useful representation of their architectures. SVGe learns a smooth graph embedding that reconstructs architectures accurately and supports efficient performance prediction and search.
2020
Weak diffraction signals can reveal geological details that conventional reflections miss. This work combines diffraction imaging with a CNN trained on synthetic examples to locate subsurface scattering points, including challenging signals in field data.
Picking geological horizons becomes more manageable when it is treated as a 3D segmentation problem. Our network learns from sparse interpreter annotations and can be refined interactively to trace complex reflection surfaces in seismic volumes.
Why do generated images leave detectable frequency fingerprints? We trace these artifacts to common upsampling operations and introduce spectral regularization that helps generators better reproduce natural-image statistics.































