Drag any node to pull the map apart. Click a line or a node to see the research tied to it.
Biomedical research faces a persistent challenge: making sense of large-scale, high-dimensional data drawn from genetics, environment, imaging, and clinical records well enough to catch a disease's mechanism early, not merely name it after the fact. Doing this honestly requires knowing when a discovery is real, when a model's predictive performance can be trusted, and when a causal claim is solid enough for an institution to act on for the public's benefit. My past and future work confronts these challenges directly, aimed at narrowing the gap between lifespan and healthspan: the years a body survives versus the years it survives well.
When Does a Discovery Reflect Real Signal, Not Dependency-Inflated Noise?
A persistent obstacle in large-scale hypothesis testing is that the theoretical null distribution routinely fails to describe the empirical distribution of test statistics once dependency enters high-dimensional biomedical data, which biases results toward false discoveries that look statistically clean.
Driven by that gap, I built an empirical null estimation method for zero-inflated discrete mixture distributions, an approach that can be read as an optimization problem minimizing the Kullback-Leibler divergence between the empirical and theoretical distributions, and extended the same protein-domain data into a rare-variant framework for tumor samples. Guided by the same concern in sparser data, I combined Bayesian estimation with empirical null modeling into a local false discovery rate method for protein domain hotspots, and, for grouped data of much higher dimension, into a correlation-adjusted simultaneous testing procedure, applied to methylation changes linked to type 2 diabetes where the dependency structure is otherwise invisible to a standard test. A related question, whether a risk difference is homogeneous across sparse count strata rather than an artifact of pooling them, needed a meta-analytic heterogeneity test built for count data rather than an adaptation of an existing one. Mentoring a graduate student on zero-inflated generalized Poisson models, characterizing the Fisher information matrix under missing data, extended this line of work into a setting especially relevant where excess zeros are common: rare disease occurrence, single-cell genomics, and microbiome studies.
The same testing machinery raises a second question I care about: can the tools built to catch false discoveries in genomics be turned toward auditing fairness in the predictive models now used in healthcare and elsewhere? Motivated by that question, I worked on Bayesian variable selection using knockoffs, a principled approach to feature selection with false discovery rate control in high-dimensional settings, as a first step toward that kind of audit.
Looking forward, I want to integrate Bayesian nonparametric methods into this line of work: Dirichlet process mixture models for empirical null estimation in heterogeneous datasets, Gaussian process approaches for functional and time-series data, and a sharper account of what a Bayesian nonparametric multiple testing procedure can guarantee. The empirical null problem's Kullback-Leibler framing, and the mutual-information reading of the prediction work described next, both point toward the same broader direction: statistical tests built on Wasserstein distances and other information-theoretic divergences that hold up when the underlying geometry of the data is not Euclidean. A recent collaboration using spectral entropy to quantify signal complexity and uncertainty in neuroimaging time series, in a regime where traditional statistical approaches struggle with a low signal-to-noise ratio, is the clearest evidence I have that this direction is not speculative.
What Does a Predictive-Performance Number Mean Once Feature Selection Is Part of the Pipeline?
A model's predictive-performance number means little once feature selection has already touched the same data used to evaluate it, and standard cross-validation treats that selection step as free, which it rarely is in high-dimensional settings.
Guided by that gap, I developed an exhaustive nested cross-validation approach that accounts for the feature selection process itself, systematically exploring all possible training-validation splits to give a more comprehensive assessment of predictive performance than traditional cross-validation, with theoretical guarantees derived from concentration inequalities in high-dimensional probability spaces that establish asymptotic control of Type I error under specific dependency structures. Read through an information-theoretic lens, this work can also be seen as optimizing the mutual information between the selected features and the response variable while controlling for overfitting. The same bias-variance tension shows up whenever features are highly correlated. Confronted by that, I worked on ridge penalization for high-dimensional testing in imaging genetics, tuning regularization parameters directly against the correlation, with results identifying genetic variants associated with brain imaging phenotypes.
It also shows up in data that is not a fixed set of variables at all. Classifying multivariate functional data, applied to fMRI recordings from patients with ADHD, required treating each patient's signal as a curve rather than a vector, and asking the same validation question in that different geometry: does a classifier trained on functional features generalize, or only fit the noise particular to one scan.
Building on this, I want to establish the asymptotic properties of exhaustive nested cross-validation in high-dimensional settings, quantify the uncertainty attached to a cross-validation-based test statistic, and develop adaptive cross-validation procedures that adjust automatically to a dataset's complexity.
What Has to Be True Before an Institution Acts on a Causal Claim?
A causal claim in biomedical research rarely gets to stay a scientific curiosity. Sooner or later it has to convince an institution, a regulator, a public health agency, a hospital, that the population it was drawn from and the study design behind it are strong enough to act on.
Motivated by that translation, I contributed to research mining and synthesizing causal relationships between diseases across the published literature to improve the use of polygenic risk scores, a step toward understanding complex disease networks and their genetic underpinnings, with implications for personalized medicine and preventive healthcare. The same question surfaces directly in public health work I have led or contributed to: as lead statistician for biosurveillance of SARS-CoV-2 in the Philippines through whole-genome sequencing, and for a proteomics-based discovery of lung cancer biomarkers and drug targets in the same country; in classifying congenital hypothyroidism at newborn screening, first with artificial neural networks and then with self-organizing maps, two different methods answering the same detection problem; and as project leader for a disability-registration framework tied to the Sustainable Development Goals. A study's adequacy is also a design question before it is a results question, which is what drove my work on power and sample size for experiments with high-frequency time-series responses. Each of these is a case of the same underlying question, asked under a different name and a different regulator: what has to be true about a causal claim, its population, and its study design before an institution is willing to act on it on the public's behalf.
The clearest future version of this question is bioequivalence and biosimilarity testing: establishing that a new therapeutic agent, biologic, chemical compound, or medical device performs comparably to an approved one, under the variability that shows up in a large patient population rather than a curated trial sample. I want to develop novel designs and statistical tests for bioequivalence studies, particularly for treatments targeting biomarkers identified through the high-dimensional methods described above, connecting the multiple-testing and cross-validation work directly to a decision a regulator has to make. The same causal-inference layer, extended with tools such as Mendelian randomization, also feeds into the biological-age work described below.
Can a Prediction of Disease Risk Carry Uncertainty a Clinician Can Act On?
Most predictive models return a single probability without a calibrated sense of how much to trust it, which is a poor basis for a decision that affects a patient's treatment.
Driven by that gap, I plan to develop a multiclass conformal prediction method for disease development and progression, giving a probabilistic assessment for an individual patient that comes with a distribution-free guarantee rather than an assumption borrowed from the model's training data. This is new work, not yet built, but it follows the same standard I hold the work above to: a claim about a patient's risk should state how confident it is, honestly.
Can Disease Progression Be Stratified Without Forcing It Into Fixed Cutoffs?
Disease progression rarely moves in a straight line, which makes it hard to say when a patient has crossed into a meaningfully different stage, or whether they will resolve on their own without ever reaching the event being studied at all.
Guided by that difficulty, I modeled clustered survival data with a cured fraction, letting a model represent that some patients in a cluster will never experience the event under study rather than forcing every survival curve toward the same eventual outcome. I extended a nonparametric version of the same clustering logic to customer survival data, a deliberately different domain that tested whether the same statistical structure, dependence within a cluster and heterogeneity in whether the event happens at all, holds outside biomedical data too.
Looking forward, I want to bring longitudinal data analysis into this same line of work: predictive models that account for fluctuations in environmental exposure and other dynamic factors over time, capturing the molecular and phenotypic variation that separates one disease subgroup, and one likely treatment response, from another, evaluated with the same nested cross-validation and false-discovery-rate control used throughout the work above so the stratification stays honest rather than overfit to a particular cohort.
Can Biological Age Close the Gap Between Lifespan and Healthspan?
All of this converges on the problem I consider the throughline of my research. Chronological age is a clock; biological age, estimated from methylation, imaging-derived brain connectivity, blood biomarkers, or telomere length, is an attempt to measure how much of the body's budget has been spent relative to that clock. I want to extend the toolkit built across the problems above, ridge-penalized regression, exhaustive nested cross-validation, false-discovery-rate control, and a causal-inference layer such as Mendelian randomization, into a full pipeline for testing whether a corrected biological-age gap predicts disease incidence and decline independently of chronological age, and whether narrowing that gap through intervention would translate into more years of healthy life rather than simply the same years of decline pushed later.
Rigor, Reproducibility, and the Responsible Use of Computational Tools
Statistical methods are powerful enough to be used well or used carelessly, and a false discovery in medical research can waste resources, sink a clinical trial, or license a harmful treatment. I hold my own work to a standard built around four commitments.
Transparency and reproducibility. Every result I produce has to be regenerable by someone who was not in the room when it was first produced, including a future version of myself. That means version-controlled code, documented pipelines, and papers that state plainly what a method can and cannot do rather than implying more than the evidence supports.
Open, reusable tools. Reproducibility is not only a property of one paper's code; it is a property of whether the next person working on a similar problem, at UK Biobank scale or otherwise, has to start from nothing. I want to build an open-source toolbox around the methods above, high-dimensional testing, nested cross-validation, and uncertainty quantification, so that the barrier to analyzing large-scale biomedical data is the size of the question, not the availability of tooling.
Clear-eyed use of computational tools, including AI. Increasingly, part of an analysis pipeline is written with the help of a large language model or a coding agent. I treat these tools as I would treat any other instrument in the lab: they can accelerate a literature search, translate a results table into readable prose, or draft code from a specification I have written myself, but they do not get to invent a claim, a citation, or a number. I am the author of every method, every commit, and every sentence that carries my name, and I have to be able to explain and defend each one without the tool in the room. Every reference is checked against its actual source before it enters a manuscript, never trusted from a model's memory.
Awareness of downstream stakes. Work on disease causality and genetics could eventually influence which patients receive which treatments, which is a responsibility that demands the highest standard of rigor and communicating findings with appropriate caution rather than overstatement, in publication and in mentorship alike.
Conclusion
My research program is organized around problems, not subfields: which discoveries are real, which performance estimates can be trusted, which causal claims are safe to act on, and how to give a patient calibrated, honest information about their own risk. Solving each has meant reaching into whichever tool the problem required, high-dimensional testing, cross-validation, functional data classification, survival analysis, causal inference, and, going forward, conformal prediction and longitudinal modeling. I am looking for a position where that combination, technical depth and a clear stake in the outcome, is what the work is judged on.
Why this page exists
The disease-network diagram on this site's research page is not a finished artifact. It is a working map of how a disease's cause becomes a clinical outcome, built one node and one edge at a time as the underlying research gets done. Two of the twelve favorite problems on the companion page, P1 and P2, live entirely inside this diagram. This page expands both into the full set of questions they actually stand for, edge by edge.
A dashed line on the diagram marks an edge the framework claims but that no paper or essay on this site documents yet. Those gaps are not a flaw in the diagram. They are, literally, the open questions below with no citation behind them yet.
P1: What is the right level to model a disease
The framework question: What is the right level at which to model how a disease actually works in the body, a gene, a pathway, a network of genes, or the whole organism, and does that answer change depending on the disease?
This single question gets asked once per edge in the diagram, because every edge is a claim about which level of resolution actually explains the next node in the chain. Eleven versions of it follow, each carrying the same three layers: the original question as it sits in the working list, a plain-language rephrase, and a paper-scoped version specific and testable enough to be an actual research question.
Genetics → Environment / Lifestyle
Documented: correlation-adjusted simultaneous testing; TXNIP methylation and type 2 diabetes in Filipinos
How much of a disease's course is already written into the genome before a single environmental exposure gets the chance to weigh in, and does that share stay fixed over a lifetime or shift as the environment does its work? In a cohort with both genotype and lifestyle-exposure data, how much variance in disease risk is explained by a polygenic risk score alone, versus the same score interacted with lifestyle covariates, after correcting for multiple comparisons across candidate exposures?
Genetics → Microbiome (gap)
Does the microbiome shape which genes get expressed, or does the genome quietly decide which microbiome a body is able to host in the first place? Does host genotype predict microbiome composition, or does adjusting for microbiome composition attenuate an existing genotype-disease association in the same dataset?
Environment / Lifestyle → Microbiome (gap)
How much of what looks like a direct environmental effect on disease is actually the environment acting through the microbiome as a middle step no one measured? Using a mediation-analysis framework, what fraction of the association between a specific environmental exposure and a disease outcome is mediated through microbiome composition, versus acting on the outcome directly?
Environment / Lifestyle → Brain Connectivity (gap)
Can an environmental exposure change brain connectivity in a way that shows up in imaging before it shows up in behavior, or does the scan only ever catch up after the fact? Does longitudinal exposure to a specific environmental variable predict a measurable change in resting-state functional connectivity prior to any detectable change in a validated behavioral outcome measure?
Environment / Lifestyle → Treatment / Intervention (gap)
Should a treatment protocol change depending on the environment a patient lives in, and if it should, why does so little outcome data actually get collected in a way that could answer that? Does a treatment-response model that includes environmental covariates outperform, under cross-validated predictive accuracy, a model trained on clinical and genetic covariates alone?
Microbiome → Brain Connectivity (gap, the gut-brain axis)
Is the gut-brain connection a real causal channel, or a correlation that keeps surviving because both organs happen to respond to the same upstream cause? After adjusting for shared upstream confounders such as diet and inflammation markers, does microbiome diversity retain a significant association with brain-connectivity measures, or does the association disappear?
Genetics → Brain Connectivity
Documented: ridge penalization in imaging genetics; predictive performance under nested cross-validation
When a genetic variant is associated with a pattern of brain connectivity, is the brain the outcome being explained, or just the most convenient place to go looking for a gene's effect? Does a genetic variant's association with a brain-connectivity phenotype replicate out-of-sample under nested cross-validation, or does it only hold within the discovery cohort?
Genetics → Treatment / Intervention
Documented: SARS-CoV-2 biosurveillance in the Philippines; clinical proteomics in lung cancer in the Philippines
How much of personalized medicine is really using the genome to predict a treatment response, versus using it to justify a decision that would have been made anyway? Does incorporating a genetic biomarker into a treatment-assignment rule change the predicted outcome distribution relative to the standard-of-care assignment rule, once evaluated on held-out data?
Demographics → Clinical Outcome
Documented: congenital hypothyroidism screening with neural networks; SDG disability-registration framework
When a demographic variable predicts an outcome, is it a cause in its own right, or a proxy standing in for everything unmeasured that happens to correlate with it? Does the association between a demographic covariate and a clinical outcome persist after adjusting for the socioeconomic and environmental variables it most plausibly proxies for?
Latent Variable → Clinical Outcome
Documented: clustered survival with a cured fraction; nonparametric clustered customer survival
If the strongest driver of an outcome is a variable nobody measured, what does it even mean to call the model that leaves it out correct? How sensitive is a survival model's estimated cure fraction to the assumed distribution of an unobserved frailty term, and does that sensitivity change the clinical conclusion drawn from it?
Treatment / Intervention → Clinical Outcome
Documented: homogeneity of risk differences under sparse counts
Once a treatment's effect is confirmed on average, how many of the individual patients it was tested on does that average actually describe? Does the estimated treatment effect differ significantly across patient subgroups defined by a specific stratifying variable, after adjusting for the multiplicity of subgroup tests performed?
Five of these eleven edges, genetics to microbiome, environment to microbiome, environment to brain connectivity, environment to treatment, and microbiome to brain connectivity, are still dashed lines on the diagram. That makes this cluster the direct bridge between the questions on this page and the paper pipeline: closing a gap in the diagram and answering one of these questions are the same act. Every edge, gap or documented, is also a candidate point of early intervention, somewhere the chain from cause to clinical outcome could in principle be interrupted before the outcome becomes years of decline.
P2: The three clocks
Chronological age is just the clock on the wall. Biological age is an attempt to measure how much of the body's budget has actually been spent relative to that clock, and brain age is one specific biological-age estimate, built from neuroimaging or connectivity data rather than blood or tissue. Brain age is not a synonym for biological age; it is one clock among several. This cluster is the most literal version of the site's whole premise: the gap between biological age and chronological age is one of the few things measurable today that stands in for the gap between lifespan and healthspan itself. A person whose biological age runs ahead of their birthday is, by this logic, someone spending more years sick relative to years lived than the calendar alone would suggest.
Typical Trajectory enters Poor Health at
Extended Healthspan enters Poor Health at
Extra years spent out of Poor Health
| Stage | Typical | Extended |
|---|
If two people share the same chronological age, what is actually being measured when a model says one of them is biologically older than the other?
Paper-scoped: After correcting for the known regression-to-the-mean bias in age-prediction residuals, does a brain-age gap estimated from resting-state functional connectivity predict incidence of a specific outcome over five years, independently of chronological age?
Is the gap between biological age and chronological age a cause of poor health, a symptom of it, or just a second ruler measuring the same thing twice?
Paper-scoped: Using longitudinal cohort data, does a widening brain-age gap precede the clinical onset of a condition, or does it only become detectable after diagnosis?
Can brain age be trusted as a stand-in for the health of the whole body, or does it only ever measure the one organ it was trained on?
Paper-scoped: How well does a brain-age gap correlate with biological-age estimates from other tissues, such as an epigenetic clock from blood, within the same individuals, and does that correlation hold across demographic subgroups?
If narrowing a biological-age gap became possible, would that buy more years of healthy life, or only delay the same amount of decline to a later date?
Paper-scoped: Does an intervention that narrows the brain-age gap in a treatment arm translate into a measurable extension of disability-free years relative to a control arm, and at what point in follow-up does that effect appear or disappear?
How this kind of research actually gets built
The four questions above are only as good as the pipeline behind them, and the pipeline is the same five-step process regardless of which clock is being measured.
- Pick the clock. Biological age is not one thing. DNA methylation, or an epigenetic clock, neuroimaging-derived brain age, composite blood-biomarker panels, and telomere length are different clocks answering different sub-questions about what is aging.
- Train an age predictor. Fit a model, often penalized regression such as ridge, on a healthy reference cohort to predict chronological age from the chosen features.
- Compute the gap, then correct it. Predicted age minus chronological age is the raw age gap, but that raw gap correlates with chronological age itself through regression dilution, so it needs a bias correction before it means anything on its own.
- Validate out of sample. The bias correction and the age gap cannot be estimated on the same data that produced them. Nested cross-validation, not a single k-fold pass, is what keeps this honest.
- Test the corrected gap against outcomes, then ask the causal question last. Once corrected, test whether the gap predicts disease incidence, decline, or mortality, ideally longitudinally and with false-discovery-rate control if many outcomes or brain regions are tested at once. An association between the age gap and an outcome is not evidence the gap is causal; that step needs a natural experiment, an actual intervention, or a causal-inference layer such as Mendelian randomization where a genetic instrument is available.
Ridge penalization, nested cross-validation, false-discovery-rate control, and causal inference are not four separate specialties sitting next to each other on this site. Here, they are the five-step pipeline one specific curiosity, the gap between how old a body is and how old it claims to be, actually needs to become an answerable question.