Thu NguyenPh.D.

Research

Selected research

Organized around scientific capabilities rather than employers: each theme spans the question being asked, why it matters, the approach taken, and the evidence produced. For the underlying record, see the publication list.

Selected research · Industry

Protein sequence design and deimmunization

Scientific question

Can protein sequences be redesigned to lower predicted immunogenicity while preserving structure and function?

Why it matters

Therapeutic proteins that provoke unwanted immune responses can lose efficacy or safety. Computational deimmunization offers a way to screen for lower-risk sequence variants before wet-lab validation.

Approach

Combined Markov chain Monte Carlo (MCMC) sequence sampling with NetMHC-predicted immunogenicity scores and inverse-folding scores from ESM-IF1 and ProteinMPNN, balancing immunogenicity reduction against structural and functional constraints. Separately, added property guidance to ADFLIP to steer deimmunized sequence design.

My contribution

  • Implemented the MCMC optimization loop combining NetMHC and inverse-folding scores under competing constraints
  • Added property guidance to ADFLIP for deimmunized sequence design

Evidence & results

70%+

NetMHC score reduction

on monomer Griffithsin

10%

TM/activity increase

reported on the same design

40%

NetMHC score reduction

on SaCas9, via ADFLIP property guidance

Methods & tools

MCMCNetMHCESM-IF1ProteinMPNNADFLIPMulti-objective optimization

Completed at Insmed Inc. described only at the level of detail present in Thu's public résumé.

Selected research · Ph.D.

Structure prediction and geometric deep learning

Scientific question

How can geometric deep-learning architectures better capture protein structure, including antibody-antigen interactions and hard-to-predict regions?

Why it matters

Accurate structure prediction underlies most downstream design work. Improving attention mechanisms and reproducing state-of-the-art pipelines makes prediction more useful for design tasks that follow it.

Approach

Reproduced AlphaFold-style monomer and multimer models in PyTorch across CUDA and ROCm environments, extended SE(3)-Transformer self-attention with additional kernels, fine-tuned AlphaFold-derived models for CDR-loop and antibody-antigen prediction, and built a peptide-design network targeting a specified binding location.

My contribution

  • Reproduced AlphaFold monomer/multimer training in PyTorch for multi-node, multi-GPU environments
  • Extended SE(3)-Transformer attention with additional self-attention kernels
  • Fine-tuned models for CDR-loop and antibody-antigen interaction prediction
  • Built a deep-learning approach for peptide sequence design targeting a specified binding location

Evidence & results

Up to 30%

precision improvement

on an n-body dataset, after adding self-attention kernels

10%

docking accuracy improvement

via Generalized Born scoring added to FTMap, measured by RMSD

Methods & tools

AlphaFoldSE(3)-TransformerPyTorchCUDAROCmDistributed trainingGeneralized Born scoringFTMap

Completed as part of Thu's doctoral research at Stony Brook University.

Selected research · Current work

Binder and ligand design

Scientific question

Can pH-sensitive binders be designed computationally to control target engagement across physiological environments?

Why it matters

pH-sensitive binding can enable more selective or tunable therapeutic behavior. This is active, ongoing work rather than a completed and validated result.

Approach

Designing pH-sensitive nanobody and ligand binders for Pdgfr-b using LigandMPNN and RFD3.

My contribution

  • Applying LigandMPNN and RFD3 to design pH-sensitive nanobody and ligand candidates

Methods & tools

LigandMPNNRFD3pH-sensitive binder design

Ongoing work at Insmed Inc. No experimental validation results are public yet, so no outcome metrics are listed for this theme.

Selected research · National lab

Structural biology and crystallography

Scientific question

How can imaging and data-processing pipelines better classify and compress macromolecular crystallography data?

Why it matters

High data-rate crystallography generates large volumes of diffraction data. Better classification and compression make large-scale structural studies more tractable and can reveal subtle structural heterogeneity.

Approach

Built a CNN-based classifier for indexable diffraction images, implemented lossy compression to reduce diffraction-image storage, and used structure-factor clustering to detect movement in binding hotspots.

My contribution

  • Developed a CNN in PyTorch to detect and classify indexable diffraction images
  • Implemented lossy compression for X-ray diffraction images to reduce storage
  • Applied structure-factor clustering to detect binding-hotspot movement, revealing small structural differences in chymotrypsinogen
  • Co-developed a technique to classify diffraction data from dynamic proteins by individual polymorph

Methods & tools

CNNPyTorchX-ray crystallographyStructure-factor analysisLossy compression

Completed at Brookhaven National Laboratory.

Selected research · Cross-cutting

Scalable scientific computing

Scientific question

How can research-scale deep learning training be made reproducible and portable across accelerator ecosystems and cloud infrastructure?

Why it matters

Scientific AI models are only useful if they can be trained reliably at scale. Portability across CUDA/ROCm and cloud batch systems reduces friction for research teams reusing these pipelines.

Approach

Built PyTorch training pipelines compatible with both CUDA and ROCm across multi-node, multi-GPU clusters, and set up folding-model workflows on Google Cloud Platform for other users via GCP Batch and Vertex AI.

My contribution

  • Made AlphaFold reproduction code compatible with both ROCm and CUDA for multi-node training
  • Set up folding models on GCP for internal users
  • Pushed training workflows to GCP Batch and Vertex AI

Methods & tools

PyTorchCUDAROCmGCP BatchVertex AIMulti-node/multi-GPU training

Spans Thu's doctoral research at Stony Brook University and current work at Insmed Inc.

Ph.D. dissertation

Doctoral research

Doctoral dissertation

Deep Learning and Its Advancements in Protein Structure Prediction

Author:
Thu Nguyen
Institution:
State University of New York at Stony Brook
Degree:
Ph.D. in Computer Science
Completed:
December 2023
ProQuest record:
2024, publication number 31297637

Research overview

Thu's doctoral work explored how deep learning can improve protein structure prediction and support protein design. Her research included reproducing and extending AlphaFold-style systems, enhancing geometric attention in SE(3)-Transformers, improving predictions in challenging structural regions, and connecting predicted structures to peptide and interaction design. The work emphasizes both model development and the engineering needed to train scientific AI across modern accelerator environments.

This overview is an original summary written for this site, not a quotation of the dissertation abstract.