Skip to content
@JinHongDu-Lab

JinHongDu-Lab

Jin-Hong Du Lab

Statistics, Causal Inference, and Reliable Machine Learning

We develop statistical and machine-learning methods for understanding complex data, making reliable decisions, and supporting scientific discovery.


About Us

Our lab led by Jin-Hong Du works at the intersection of statistics, machine learning, and data-driven science. Our research combines rigorous statistical theory with modern computational methods to address problems involving causality, interpretability, distribution shifts, and complex structured data.

Research Areas

Our current research includes:

  • Causal inference Identification, semiparametric inference, treatment-effect estimation, transportability, and causal learning under weak assumptions.

  • Reliable and interpretable machine learning Feature importance, uncertainty quantification, distribution-shift robustness, and principled evaluation of machine-learning systems.

  • High-dimensional and structured data Statistical methods for networks, time series, latent-factor models, functional data, and other complex dependent data.

  • AI for scientific discovery Machine learning for single-cell genomics, perturbation modeling, biomedical research, and other scientific applications.

  • Foundation models and intelligent systems Statistical foundations for reasoning, causal representation, out-of-distribution generalization, and verifiable evaluation of foundation models and agents.

What You Will Find Here

This organization hosts:

  • Research software and reproducible implementations
  • Code accompanying our papers and preprints
  • Simulation and benchmarking frameworks
  • Tutorials, examples, and research resources
  • Collaborative projects developed by lab members

Repositories will include documentation, reproducibility instructions, and licensing information whenever possible.

Research Principles

We aim to build methods that are:

  • Statistically principled
  • Computationally practical
  • Transparent and reproducible
  • Robust to real-world complexity
  • Relevant to substantive scientific questions

Collaboration

We welcome collaborations across statistics, machine learning, artificial intelligence, genomics, and data-intensive scientific disciplines.

For research information, publications, and updates, visit Jin-Hong Du’s website.


Theory with purpose. Methods that survive contact with data.

Popular repositories Loading

  1. scVAEIT scVAEIT Public

    Variational autoencoder for single-cell integration and transfer learning.

    Jupyter Notebook 10

  2. causarray causarray Public

    causarray is a Python module for simultaneous causal inference with an array of outcomes.

    Jupyter Notebook 7

  3. crispyx crispyx Public

    crispyx is a Python package for scalable CRISPR and Perturb-seq screen analysis on disk-backed AnnData files. It provides streaming quality control, normalization, pseudobulk aggregation, and diffe…

    Python 3

  4. ensemble-cross-validation ensemble-cross-validation Public

    Cross-validation methods designed for ensemble learning

    Python 1

  5. FDFI FDFI Public

    Disentangled feature importance

    Python

  6. .github .github Public

Repositories

Showing 6 of 6 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…