Skip to content

Medical imaging · implicit neural representations · deformable registration

Anatomy as a continuous function, not a grid of numbers.

I am finishing an M.S. in Computer Vision at the University of Central Florida, where I work on how to represent the deformation between two human bodies. My first author paper, QueryPhi, replaces the fixed voxel lattice that almost every method uses to store that deformation with a field you can evaluate at any coordinate.

It reaches the highest mean Dice of all ten methods on the IXI brain benchmark using a fraction of the parameters, keeps the warp near fold free, and carries the best boundary accuracy to a cohort it never trained on.

Applying for PhD positions starting 2027Orlando, Florida
IXI atlas benchmark
1 of 10
highest mean Dice, 30 structures
Mean Dice
0.7726
95% CI 0.769 to 0.776
Model size
5.51M
8.5 times smaller than the transformers
Folded voxels
0.0001%
near fold free, measured

What I work with

Deformable and diffeomorphic registrationLDDMM and B-spline parameterisationsFoundation model adaptation: SAM, LoRANIfTI and DICOM handlingDifferential geometry of deformationsPyTorchMONAIscikit-learnOptuna hyperparameter searchWeights and BiasesTensorBoardMulti GPU and distributed trainingMixed precision, gradient checkpointingLoRA and parameter efficient fine tuningVision transformers, Swin, nnFormer, CoTrPaired Wilcoxon signed rank testsWilson intervals for ratesAblation and attribution designPythonMATLAB
C and C++ fundamentalsTypeScript and JavaScriptBash and shell automationLaTeXGit and GitHubNumPy, SciPy, PandasMatplotlib, Seaborn, PlotlyOpenCV, scikit-image, PillowAdobe Lightroom, Photoshop, Premiere ProFFmpegLinuxHPC and SLURM job arraysGPU accounting and fair share budgetingDockerDual H100 and A100 nodesReproducible experiment trackingRegression checks on scientific resultsNext.js, React and TailwindPyQt, FFmpeg, ISO 13485 documentation

The question

Putting two anatomies into correspondence

Given two scans of different people, deformable registration finds the smooth, invertible map that carries one onto the other. What that map has to do in order to succeed is precisely the difference between those two bodies. Atlases, morphometry, longitudinal studies and label transfer all rest on it.

The convention is to store that map as a displacement vector at every voxel of a fixed grid. QueryPhi stores it as a function instead, so the deformation has a value at every coordinate rather than only at the grid nodes, and its derivatives come out analytically.

Moving image, fixed image and the warped result

Two different people, and the smooth invertible map that carries one onto the other. What the map has to do in order to succeed is the difference between them.

NeurIPS 2026 · under review · first author

QueryPhi: A Compact Pair-Conditioned Continuous-Query Model for Diffeomorphic 3D Registration

A registration model that stores the deformation as a function you can evaluate anywhere, rather than as numbers parked on a fixed voxel grid.

0.7726

IXI atlas Dice

first of ten, 95% CI 0.769 to 0.776

0.7534

Nearest method

TransMorph-bspl, p = 1.5e-8

5.51M

Parameters

against 46.8M for the transformers

0.0001%

Folded voxels

near fold free, empirical

1.502 mm

OASIS zero shot

best HD95 of nine methods

1.77

Average rank

Friedman and Nemenyi, ten methods

A modulated SIREN decodes the velocity field as a continuous function of coordinates, over a coarse to fine correlation pyramid. Scaling and squaring integrates that velocity into the deformation. The whole model is 5.51 million parameters and trains without ever seeing a segmentation label.

Implicit neural representationsDiffeomorphic registrationSIRENStationary velocity fieldsPyTorchH100
The full account
QueryPhi architecture: encoder, correlation pyramid, modulated SIREN, scaling and squaring

Encode, correspond, query, integrate, warp. The SIREN in the middle is the part that makes the field a function rather than an array.

Results at a glance

All 13 figures
Accuracy against topology preservation. Bubble area is parameter count. QueryPhi sits far right in the near fold free band while being the smallest model near the top.
Qualitative comparison across all ten methods on the same case.
Critical difference diagram from a Friedman test with Nemenyi post hoc comparison over ten methods and 115 paired cases.
Per case Dice distributions rather than means alone. QueryPhi is both the highest and the tightest.

4 items. Drag, scroll, or use the arrows.

Where this goes next

Input resolution, and anisotropic acquisition

The continuous query in QueryPhi concerns the decoder output. The encoder and the correlation pyramid still run at the resolution the model was trained on, so the paper is explicit that robustness to different input voxel spacings, and to anisotropic acquisition, is not something it demonstrates.

That gap is the most interesting thing about the result. A continuous field should earn its keep exactly where sampling is the binding constraint, and on 1 mm isotropic brain MRI sampling is generous. Most clinical MRI outside neuroimaging is not: an abdominal sequence might be 1.5 mm in plane and 7 mm between slices, close to five to one.

So the work in progress is a controlled study on a benchmark that contains both, with anisotropy as a dial inside one dataset rather than a difference between datasets, the endpoint pre-specified as topology rather than accuracy, and every method held to the same protocol. Early search level runs are encouraging and are nowhere near the fidelity needed to make the claim, which is why the paper lists this as a limitation rather than a finding.

The caveats in full
Re-querying the trained field across output resolutions

The field can be re-queried at any output resolution once trained. Whether that also holds for anisotropic input sampling is the open question, and the paper says so rather than implying otherwise.

CVPR 2026, PVUW Workshop · co-author

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

A benchmark asking whether vision language models actually understand interactions in video, or recognise objects and guess the rest.

VISTA compared against existing spatio-temporal benchmarks

Across 11 state of the art models, breaking aggregate performance down along the taxonomy revealed pronounced spatio-temporal biases that traditional metrics obscure completely. Models that looked comparable in aggregate turned out to behave very differently once the axes were separated.

VISTA is the first large scale interaction aware diagnostic benchmark for spatio-temporal understanding in vision language models.

Projects

Things I built along the way

Course work, challenge entries and implementations written from scratch. Each one is here because it taught me something that shows up in the work above.

The geometry of generalization output
MAP 6197 Mathematical Introduction to Deep Learning2025

The geometry of generalization

Four training runs on paediatric brain tumor MRI, held identical except for the loss and the optimiser, then read through the curvature of the loss surface rather than the leaderboard.

PyTorchResidual U-NetPyHessianSAM and Lion optimisers
01 / 06

Writing

Notes on the work

Longer pieces on how the method came together, what the measurements said, and the experimental habits I picked up on the way.

How I got here

I came for the mind and stayed for the body carrying it

I went toward brain imaging for reasons that were more philosophical than technical, and the interest moved once I arrived. What holds me now is anatomy, and the geometry and optimisation sitting underneath the methods that measure it.

The habit I cannot put down is wanting to know which part of a result is doing the work. Most of what I have built since is an attempt to make that question answerable rather than arguable.

More about me

Hover for colour, click to enlarge.

Currently

Where things stand

  1. Teaching Assistant, CAP 4453 Robot Vision

    Fall 2026

    University of Central Florida

  2. Research Assistant, Medical Image Analysis

    January 2026 to present

    AI MIND Lab, University of Central Florida

  3. Research Assistant, Vision Language Evaluation

    August 2025 to January 2026

    Center for Research in Computer Vision, UCF

Full history

Applying for PhD positions starting 2027

Medical imaging, continuous and physically constrained representations, uncertainty, and reconstruction from incomplete measurements. Happy to share the manuscript.

Get in touch