Matthew Vaishnav
Applied ML / Computational Pathology Research Engineer
Independent neural-network research in whole-slide histopathology, scanner and acquisition robustness, representation auditing, multiple-instance learning, and reproducible ML systems.

Work
I am an independent computational pathology engineer and applied machine-learning researcher. I build controlled experiments and reproducible ML systems for whole-slide modeling, pathology foundation-model features, scanner and site robustness, representation audits, simulated federated learning, and fail-closed research infrastructure.
My primary research line is Paired-Acquisition Neural Factorization. Using multiple scans of the same underlying tissue, I test whether frozen pathology embeddings can be separated into a tissue-oriented representation with substantially reduced linearly recoverable scanner identity and an acquisition branch that retains scanner information, while preserving descriptive tissue-category structure and same-region retrieval under the tested protocols.
Current Research
1. Paired-Acquisition Neural Factorization
Corrected, fold-aware SCORPION evaluation across 48 human H&E slides, five scanners, and DINOv2, Phikon, and ResNet50 feature families, together with a 175-fit capacity-matched ablation campaign.
2. External multi-scanner validation
Independent canine squamous-cell carcinoma validation using biological-sample-blocked folds, a corrected fixed five-category audit, and a completed 450-cell dimensionality × cross-covariance factorial.
3. Prospective linear baseline comparison
Preregistered comparison against paired affine and orthogonal-Procrustes controls to separate the value of neural factorization from simpler harmonization. No comparative result is claimed before execution and promotion.
4. Pair-repeat allocation
Matched-budget experiments testing unique biological pair diversity against repeated exposure to the same anchors.
5. CAMELYON17 center-subspace projection
Mechanism-focused work on attenuating source-center information while auditing tumor signal in frozen pathology representations.
6. Whole-slide multiple-instance learning
PANDA slide-level modeling with mean pooling, gated AttentionMIL, and a repaired TransnnMIL implementation. Historical fusion scores are retained only as records; matched reruns are required for new architecture claims.
7. Research reliability infrastructure
Immutable provenance, artifact hashing, corruption tests, resumable factorial runs, fail-closed validators, preregistered analyses, and dedicated GitHub Actions gates.
Selected Evidence
- SCORPION study scale
- 48 / 480 / 5
- Slides / aligned regions / scanners
- SCORPION scanner probe
- 0.7825 → 0.3989
- Reduced linear scanner recoverability
- Capacity-matched campaign
- 175 / 175
- Registered fits validated
- Canine SCC factorial
- 450 / 450
- No universal operating point found
- PatchCamelyon test
- 0.9394 AUC
- 0.8526 accuracy on one official split
- PANDA readable features
- 10,611
- Verified slide-level feature vectors
Claim Boundary
Research-only. Not clinically validated. Not diagnostic software. Not intended for clinical deployment or patient-care use. The current paired-acquisition evidence supports partial structured separation under the tested conditions: substantially lower linearly recoverable scanner identity in the tissue-oriented branch, strong scanner information in the acquisition branch, and preserved descriptive tissue-category structure and same-region retrieval. It does not establish pure biological factors, complete scanner invariance, disease biology, clinical utility, or deployment readiness.
Read the authoritative claim boundaryBio
I ♥
Matrix multiplication, backpropagation, gradient descent, optimization landscapes, attention mechanisms, convolutional inductive biases, embedding geometry, latent-space factorization, feature disentanglement, multiple-instance learning, pathological failure modes, and figuring out what neural networks actually encode.
Inspired by Takuya Matsuyama's homepage