Geongyu Lee · 이건규
Geongyu Lee · 이건규 AI Researcher · Seoul

Geongyu Lee이건규

Computational Pathology / Pathology Foundation Models / Multi-Omics

Five years building deep learning for medical imaging and drug discovery: from whole-slide-image pipelines and MFDS-track validation to pathology foundation models, now integrating proteomics and multi-omics.

5+
Years · deep learning R&D
5
Peer-reviewed journals · 2 (co-)first
G2L
Pathology FM · AAAI 2026 Workshop oral
2nd
KPI Challenge 2024 · MICCAI
selected results
0.85
2-year Prostate BCR AUC · internal test
0.76
Breast WSI mIoU · 80% coverage
91.2%
High-risk classification accuracy
≥0.65
LDO CV Pearson · held-out drugs
00

Profile

I build deep learning systems for computational pathology and drug discovery. First-author work on H&E-based breast-cancer recurrence in Scientific Reports, a co-first review in Prostate International, and a contribution to large-scale pathology foundation models (G2L, AAAI 2026 workshop, oral).

At OMIXAI I'm extending histopathology AI into proteomics and multi-omics, including a Korea–US–Japan precision-oncology collaboration on Pan-Sarcoma proteogenomics. A recurring thread is predicting molecular state from H&E alone (gene expression, spatial proteomics, receptor status) and testing whether those predictions can actually stand in for a molecular assay. Earlier, at Deep Bio, I led WSI pipeline development and pre-deployment model validation for an MFDS-track product.

Research Interests

  • Pathology foundation models
  • Histology-to-molecular prediction (gene expression, spatial proteomics)
  • Proteomics & multi-omics × histopathology
  • Virtual-cell & perturbation prediction
  • Clinical decision support (human & veterinary)
  • Calibration & uncertainty in clinical AI
  • Self-/weakly-supervised learning

Leadership at a glance

  • Led / co-led 4 research lines: WSI pipeline, prostate metastasis & recurrence risk, multi-omics FM R&D, Virtual Cell Challenge
  • First-authored a 2-hospital multi-center cohort study (NCC · KU Guro) · built a 500+ slide WSI pipeline
  • Researcher on 3 national / international programs · 3-country (KR · US · JP) precision-oncology collaboration
  • Mentored 30+ early-career AI/data professionals 1:1 · guided 8+ project teams across 3 cohorts · Codeit mentor
01

Publications

† (co-)first author · Google Scholar ↗
01 †
Grad-CAM heatmaps over H&E breast tissue patches Scientific Reports · 2025 · First author

Recurrence risk in early-stage breast cancer from H&E whole-slide images

Multi-center retrospective study (n=125; National Cancer Center + Korea University Guro) showing histopathology-only deep learning can approximate genomic-assay-style risk stratification.

Sensitivity0.86 / 0.75 / 0.53Specificity0.82 / 0.80 / 0.97
Paper ↗
02
G2L knowledge-distillation framework overview AAAI 2026 Workshop (W3PHIAI) · Oral

G2L: giga-scale to cancer-specific pathology foundation models

Distilling giga-scale pathology foundation models into compact, cancer-specific encoders via knowledge distillation, preserving downstream performance at lower compute.

Knowledge distillation·Self-supervised
arXiv ↗
03
KPIs 2024 patch-level glomeruli: PAS image, ground-truth mask and overlay Medical Image Analysis · 2026

KPIs 2024 Challenge: glomerular segmentation from patch- to slide-level

The official challenge report (Medical Image Analysis, vol. 114, 104234) for kidney glomeruli segmentation, 2nd place in the whole-slide track, tackling generalization from patch-level training to slide-level inference.

Whole-slide track2nd place·MICCAI 2024
Paper ↗
04 †
AI-based digital pathology workflow from specimen slide to diagnostic report Prostate International · 2025 · Co-first

AI-driven digital pathology in urological cancers

A review of current trends and future directions for AI in urological cancer pathology, spanning clinical, technical, and regulatory perspectives.

ReviewDigital pathology
Paper ↗
05
MurSS multi-resolution selective segmentation architecture Bioengineering · 2024

MurSS: multi-resolution selective segmentation for breast cancer

Ingests multiple resolutions simultaneously and abstains on unreliable regions via a confidence-based reject loss, built for noisy histopathology labels.

mIoU0.728 · 95% coverage·0.760 · 80% coverage
Paper ↗
06
Tumor patch risk-score heatmap with high- and low-risk patch regions Preprint · 2026

Spatial proteomics guided by H&E-based AI: recurrence-risk niches in TNBC

Using H&E-based AI risk scores to guide spatial proteomics across 156 triple-negative breast cancer patients. High-risk niches showed proliferative programs and low-risk niches immune activation, yielding a 13-protein signature that improved prognosis when combined with the image-based risk score.

H&E risk modelAUC 0.77 / C-index 0.77· independent testH&E + 13-protein compositeOOB C-index 0.679 → 0.739
arXiv ↗
07 †
MoSPR: patches to morpho-spatial macrostates, then low-rank regression to gene expression Preprint · 2026 · Co-first

MoSPR: histology-to-gene expression prediction with morpho-spatial macrostates

A linear framework that groups WSI patches into morphological microstates, discovers spatial macrostates from their adjacency, and maps them through low-rank molecular programs to gene expression. Outperforms 14 baselines across three TCGA cancers, stays strong with less training data, and its linear form attributes each prediction to global and region-specific programs.

Gene-wise PCCBRCA 0.413 · KIRC 0.334 · LUAD 0.358· best of 15 methods
08
Multi-section whole-slide sampling strategies for prostate BCR Preprint · 2026

Multi-section WSI analysis for biochemical recurrence in prostate cancer

Efficient AI-driven analysis across multiple tissue sections to predict biochemical recurrence, extending single-slide inference to full multi-section specimens.

Multi-section WSIProstate BCR
arXiv ↗
09
Encoder embedding-space illustration for supervised contrastive learning IEEE Access · 2021

Supervised contrastive embedding for medical image segmentation

Region-annotation-level pixel-wise contrastive loss to stabilize segmentation under scarce and unreliable medical-image labels; the line of work behind my M.S. thesis.

Contrastive learning·Domain robustness
IEEE ↗
02

Conferences

† first author · originals in the GitHub archive ↗
01 †
BIOINFO/GIW ISCB-Asia 2026 · Accepted poster · First & presenting author

Predictability is not substitutability: a cost-of-substitution framework for H&E-based molecular prediction across 5 cancers

Seoul · Nov 17–20, 2026. High accuracy does not make an H&E model a substitute for a molecular test. We define substitution cost by translating each misclassification into a deviation from a prespecified treatment routing, and apply one pre-registered protocol (UNI + CLAM-SB, site-disjoint hold-outs, label-shuffle controls) across five cancers. Of ~15 endpoints only HPV status in head and neck met the confirmation criterion; most were reported as undecided rather than negative. With Pseudo Lab.

ConfirmedHNSC HPV · AUROC 0.959Excludedlung subtype · V(site, label) = 1.0
02
BIOINFO/GIW ISCB-Asia 2026 · Accepted oral · Co-author

A reliability map for per-gene multiome RNA velocity parameters in single-cell kinetics

Seoul · Nov 17–20, 2026. Are the per-gene parameters that multiome RNA-velocity methods emit (transcription rate α, degradation rate γ, chromatin-to-transcription lag) actually reliable? Five velocity arms tested on four axes: cross-method reproducibility, an ATAC-shuffle control, replication in five external multiomes, and anchoring to measured synthesis and degradation rates. Only α reproduced across methods; γ and the lag need orthogonal validation.

α reproducibilityρ = 0.88γρ ≈ −0.1pre-registered 6 / 6 passed
03 †
Grad-CAM over H&E for protein-receptor status prediction AACR Annual Meeting 2023 · Poster · First author

Predicting protein receptor status from H&E-stained images in breast cancer

Abstract #5404 · April 2, 2023. ER, PR and HER2 status from H&E WSIs alone, without IHC. TCGA-BRCA (728 cases), a multi-task model with a shared morphology encoder and per-receptor heads, plus confidence-based patch selection. With Korea University Guro Hospital.

Slide-level accuracyER 74.6 · PR 66.0 · HER2 76.6
Poster ↗
04 †
Patch-level classification and majority-voting pipeline for recurrence risk AACR Annual Meeting 2022 · Poster · First author

Recurrence risk prediction based on automatic histologic analysis of breast cancer using whole slide images

New Orleans · April 12, 2022. Can H&E WSI analysis approximate the 21-gene recurrence score? 125 hormone-positive, node-negative, HER2-negative cases; patch-level CNN with majority voting into low / intermediate / high risk, nested cross-validation. No low-risk case was called high-risk. The work that grew into the Scientific Reports (2025) paper.

Accuracy L / I / H0.832 / 0.776 / 0.912Sensitivity0.857 / 0.746 / 0.529
Slides ↗
05
5-fold C-index results for pancreatic adenocarcinoma survival model AACR Annual Meeting 2022 · Poster · Co-author

A deep learning based pancreatic adenocarcinoma survival prediction model applicable to adenocarcinoma of other organs

New Orleans · April 2022. Survival risk from pancreatic adenocarcinoma H&E, with a general adenocarcinoma feature extractor (GAFE) to test transfer to other organs. TCGA pancreas (179 cases), 5-fold CV; transfer evaluated on rectum (READ) and breast (BRCA).

C-index0.726READ transfer0.694BRCA transfer0.571
Slides ↗
06
Kaplan-Meier survival curves by AI risk score USCAP Annual Meeting 2022 · Poster · Co-author

Breast cancer survival analysis through the extracted feature from the prostate diagnosis model

March 18, 2022. Breast-cancer death risk from H&E WSIs using features from a pre-trained prostate diagnosis model instead of a general-image backbone. TCGA-BRCA (980 WSIs), attention-guided multiple-instance learning, 5-fold CV; Kaplan-Meier separation of high- vs low-risk groups.

C-index0.655log-rankp = 0.0011
Poster ↗
07 †
Class-activation maps for breast histologic grading USCAP Annual Meeting 2022 · Poster · First author

Automatic histological grading of breast cancer resection tissue

March 17, 2022. Consistent histology grading (tubule formation, nuclear grade, mitoses) to reduce inter-observer variability. 125 H&E WSIs graded and annotated by a pathologist; tumor patches classified into grade 1 / 2 / 3, with class-activation maps for inspection. With Korea University Medicine.

Patch accuracy G1 / G2 / G30.958 / 0.922 / 0.956
Poster ↗
08 †
U-Net with feature-wise contrastive loss on encoder embeddings KIIE Fall Conference 2020 · Oral · First author

Utilizing a contrastive loss to improve segmentation model performance

With Prof. Sangheum Hwang, Dept. of Data Science, SeoulTech. A feature-wise contrastive loss on the encoder bottleneck pulls same-class embeddings together and pushes different classes apart. Gains for both U-Net and UNet++ on CT lung (Kaggle) and liver (LiTS2017) segmentation, largest on distance-based boundary metrics. The work behind the M.S. thesis and the IEEE Access (2021) paper.

Liver DSC0.912 → 0.923Liver ASD0.533 → 0.365
Slides ↗
03

Projects

medical AI · multi-omics · drug discoverycode ↗ github.com/Geongyu
Multi-omics · OMIXAI

Proteomics drug-response prediction

Self-supervised proteomic representation learning with LoRA / PEFT fine-tuning; also cell-line IC50 prediction. Validated under a leave-drug-out protocol on held-out compounds.

LDO CV Pearson≥ 0.65· held-out drugs·PEFT
Precision oncology · OMIXAI

Veterinary oncology CDSS & Virtual Cell

Canine oncology drug-recommendation algorithm toward a clinical decision support system, plus co-leading the Virtual Cell Challenge, predicting CRISPR-knockdown response in pluripotent stem cells.

Top-k accuracy≥ 70%· canine oncology cohort·CRISPR perturbation
FlyDiscovery workbench: PARP1 × niraparib structure, docking and affinity evidence
Side project · NVIDIA Hackathon 2026

FlyGate: evidence-first agent from drug discovery to pharmacovigilance

NVIDIA Korea Agentic AI Hackathon 2026 entry (team of 5) linking pre-market target binding (FlyDiscovery) and post-market adverse-event surveillance (FlyVigilance) in one agent. I built FlyDiscovery, which scores candidates from structure prediction to affinity with BioNeMo NIM (OpenFold3 · DiffDock · Boltz-2), and its discovery-workbench UI.

OpenFold3 RMSD1.0 ÅBoltz-2 affinityρ 0.767live ↗code ↗
H&E WSI tiling to slide-level risk prediction pipeline
Digital pathology · Deep Bio

Oncotype DX recurrence prediction

Predicting breast-cancer prognosis from H&E WSIs only, with no other clinical inputs. Patch-wise classification with confidence-aware selection and majority voting into slide-level GHI-RS risk groups.

High-risk classification acc91.2%Overall~87.3%· internal validation
Additional R&D · Deep Bio

ADMET / DTI prediction and OOD-driven active learning

ADMET prediction and drug–target interaction modeling for early drug-property profiling, plus an active-learning framework that defines out-of-distribution pathology samples to make labeling more efficient. Internal R&D, no public results.

ADMET·DTI·OOD detection·Active learning
Brain CT hemorrhage segmentation and class-activation maps
Medical imaging · SK / Ajou Univ. Hospital

Brain-hemorrhage detection on CT

2D and 3D segmentation of hemorrhage regions plus slice-level classification. Validated domain generalization on external public data and used class-activation maps to audit classification reliability (U-Net vs DeepLab).

2D + 3D seg·Domain generalizationcode ↗

Earlier ML engineering

2019–2020 · industrial vision, tabular data, data engineering, segmentation robustness · 4 projects
Machine-vision screw localization with circle detection
Computer vision · Fronttec

Screw defect classification

Defect detection on a manufacturing line: 969 images across 5 classes, under a hard 0.3s/image budget. Used machine vision (not DL) to localize the screw and strip background before classification, which lifted accuracy.

5 classes·<0.3s / image
Tabular ML · Zigbang

Apartment price prediction

Transaction-price modeling for Seoul/Busan from 1.6M records. Re-matched addresses by postal code to fix mislabeled districts, and enriched with public-API features (nearby subway, schools) missing from the raw data.

1.6M records·Metric: RMSE
Data engineering · SeoulTech

Music-chart crawling pipeline & database

Daily Top-100 charts from 6 streaming platforms (Melon, Genie, VIBE, YouTube Music, Apple Music, Bugs) crawled at a fixed time, merged on track + artist, and loaded into a relational DB. Schema designed for chart-in/out sparsity (nullable rank columns) and multilingual titles (nvarchar).

6 platforms·Daily crawl·Schema designcode ↗
Computer vision · SeoulTech

Adversarial robustness for semantic segmentation

U-Net segmentation on PASCAL VOC (20 classes, 1,928 images, 4,203 objects) with adversarial training, label smoothing and cut-out applied alone and in combination, comparing the robustness–accuracy trade-off; the starting point of the robustness-first view carried into medical AI.

Adversarial training·Label smoothing·Cut-outcode ↗
04

Experience

OMIXAIFeb 2025 – Present
Seoul · fmr. RadiSen

AI Researcher

Multi-omics · Pathology foundation models
  • Lead multi-omics foundation-model R&D integrating proteomics with H&E pathology for drug-response prediction; co-authored the G2L pathology FM.
  • Co-authored spatial-proteomics work using H&E-based AI risk scores to pinpoint recurrence-risk niches in triple-negative breast cancer (preprint, 2026).
  • Built proteomics DRP (Leave-Drug-Out Pearson ≥ 0.65) and a canine veterinary oncology drug-recommendation algorithm (top-k ≥ 70%).
  • Advanced ADMET prediction and self-supervised proteomic representation learning with LoRA / PEFT.
  • Co-led the Virtual Cell Challenge: CRISPR-knockdown response in pluripotent stem cells.
Deep BioMar 2021 – Jan 2025
Seoul · Guro

AI Researcher

Computational pathology · MFDS-track model validation
  • First-authored the multi-center breast-cancer recurrence study (Scientific Reports 2025); built a scalable WSI pipeline handling 500+ slides.
  • Developed lymph-node metastasis detection and the KPIs 2024 glomeruli segmentation model (2nd place, whole-slide track, MICCAI); challenge results published in Medical Image Analysis (2026).
  • Led prostate metastasis & recurrence-risk projects; presented at AACR (2022, 2023) and USCAP (2022).
  • Led pre-deployment model validation and contributed to regulatory documentation for an MFDS-track medical AI product; built internal server-automation tooling (Docker infra).
Nuricon2021
Pangyo

Intern

Computer vision
  • Built a parking-lot fire-detection AI system.
05

Education

SeoulTech2019 – 2021
Seoul

M.S. in Data Science

Advisor: Prof. Sangheum Hwang
  • Thesis: contrastive loss for improving segmentation under uncertain medical-image labels.
06

Leadership & Community

outside of core roles · community research · reviewing · mentoring
Pseudo Lab2026 · Season 12
Open research community

AutoBioX: AI Agents for End-to-End Bio Research

Research member (Runner) · 16 weeks · Completed
  • Completed a 16-week season project on AI agents for biological research as a research member; a nine-person team with weekly Friday sessions.
  • Two research outputs from the project were written up and accepted at BIOINFO/GIW ISCB-Asia 2026 (poster, first and presenting author; oral, co-author), then submitted to the ML4H 2026 Findings track (under review).
ML4H2026
Academic service

Reviewer

Machine Learning for Health Symposium
  • Peer reviewer for ML4H 2026 submissions.
Codeit2025 – Present
Part-time

AI Career & Project Mentor

Mentoring · Project guidance · Career coaching
  • Mentored 30+ early-career AI and data professionals through structured one-on-one résumé, portfolio, and career-development programs.
  • Supported 8+ project teams (4–6 members each) across 3 cohorts through project scoping, methodology design, experimentation, troubleshooting, and result interpretation.
  • Reviewed assignments and project deliverables, providing actionable feedback on technical quality, problem formulation, and presentation.
  • Advised mentees on career direction, project positioning, and communicating their technical contributions to recruiters and interviewers.
07

Toolkit & Funded Work

Deep Learning

PyTorchHuggingFacePEFT / LoRAAccelerateDDP / FSDP

Pathology Stack

OpenSlideQuPathCLAM / weakly-supervised MILUNI · CONCH · VirchowWSI tiling

Bio / Omics

scanpy · AnnDataSingle-cell perturbationProteomic representationADMET

MLOps & Languages

PythonR · SQL · BashDocker · SlurmW&B · MLflowFastAPI · Git

Pan-Sarcoma Proteogenomic Profiling for Precision Oncology

Korea–US–Japan · Kyung Hee Univ. Medical Center
Apr 2025 – Dec 2028 · Total program budget: KRW 400M · Participating Researcher

Virtual-Cell CDSS for Veterinary Oncology

IPET / Ministry of Agriculture · National R&D Program
2026 – 2030 (awarded) · Total program budget: KRW 3B+ · Participating Researcher

General-Purpose AI for Cancer Pathology Diagnosis

Lead: Deep Bio · National R&D Program
Apr 2021 – Dec 2025 · Total program budget: KRW 2.375B · Participating Researcher
2nd

KPI Challenge 2024 · Whole-Slide Track
Glomeruli segmentation · held at MICCAI 2024 · published in Med. Image Anal. 2026

Open to collaboration
Clinically grounded AI,
pathology to multi-omics.

Foundation models, multi-omics, or translational digital pathology: I'd like to hear about it.