About & Resume
Biography
During my Ph.D., I immersed myself in advanced scientific research, tackling complex biological problems with a focus on data-driven discovery. I developed automated data analysis pipelines where there were none, serving as the first computational (dry-lab) graduate in my lab's 30+ year history. My work centered on developing novel machine learning models to analyze intricate datasets and uncover patterns with significant biological implications. The infrastructure and knowledge base I established continued to power high-impact publications long after my graduation.
As a Senior Bioinformatician at the Hospital for Sick Children (SickKids), I design predictive algorithms and machine learning frameworks tailored for high accuracy and computational efficiency. My hands-on experience spans building diagnostic pipelines that pinpoint previously undiagnosable rare genetic disorders to automating clinical injury surveillance systems.
Beyond modeling, I build and optimize end-to-end data pipelines using workflow engines like Snakemake and WDL, enabling scalable processing of large-scale genomic and clinical datasets. One of my key achievements includes developing an NLP framework for processing unstructured clinical notes with 99%+ accuracy, reducing manual effort by over 80%.
Curriculum Vitae
Work Experience
Senior Bioinformatician
- Lead programming and analytics teams in developing scalable multi-omics and clinical data analysis pipelines.
- Architect end-to-end data infrastructures for massive clinical research datasets across genomics, NLP, and computer vision.
- Lead data analysis efforts in local, national, and international collaborations with researchers, clinicians, and pharmaceutical leaders.
- Develop state-of-the-art NLP transformer workflows achieving 99%+ accuracy on unstructured emergency department clinical notes.
- Host institution-wide and city-wide seminars covering statistical modeling, machine learning, data visualization, and omics analysis.
- Mentor junior bioinformaticians, data scientists, and clinical researchers in statistical testing and pipeline execution.
Bioinformatician
- Utilized multiple programming paradigms to automate data processing, statistical analysis, and report generation.
- Trained supervised and unsupervised machine learning models on diverse datasets (genomics, surveys, clinical notes, images).
- Consulted principal investigators on study design, experimental setup, and grant applications.
- Created interactive R/Shiny web applications for clinicians to explore and visualize complex patient cohorts.
Co-Investigator
- Serve as lead quantitative data expert, assisting with study design, grant writing, and quantitative protocol development.
- Responsible for statistical modeling of patient-reported outcomes (PROs) and mixed-methods survey integration.
Molecular Data Management Specialist
- Managed large-scale genomics repositories (microarray, NGS) for multi-site research consortiums.
- Developed open-source R and Python packages for automated ETL workflows and database querying.
- Collaborated with web developers to design user-friendly data visualization interfaces.
Ph.D. Thesis Researcher
- Established the core computational data infrastructure for 100+ Next-Generation Sequencing datasets.
- Discovered novel insights into mRNA decay pathways and translation fidelity using high-throughput sequencing.
- Engineered high-speed, memory-efficient bioinformatics algorithms to process large-scale transcriptomic data.
- Published findings in peer-reviewed journals including RNA, eLife, and Methods in Enzymology.
Statistician
- Analyzed longitudinal survey data and wearable sensor data in multi-center clinical trials on physician burnout.
- Co-authored studies in peer-reviewed medical education journals.
Education
Ph.D. in Bioinformatics & Molecular Biology
Thesis: mRNA Decay Pathways Use Translation Fidelity and Competing Decapping Complexes for Substrate Selection
B.Sc. in Economics & Molecular Biology
Karen T. Romer Undergraduate Teaching and Research Award (2006–2007)
Core Technical Skills
Programming & Frameworks
Python: PyTorch, pandas, NumPy, SciPy, scikit-learn, OpenCV, Hugging Face, spaCy, SQLAlchemy
R: Bioconductor, Tidyverse, Shiny, data.table
Web & DB: HTML/CSS, SQL (PostgreSQL, SQLite), REST APIs
HPC & Infrastructure
Workflow Engines: Snakemake, WDL
Containers & Cloud: Docker, Git, Linux / Bash
Schedulers: Slurm, LSF, Moab
Multi-Omics & Clinical Data
Genomics: WGS, WES (SNV/SV calling & annotation)
Transcriptomics: Bulk RNA-Seq, scRNA-Seq, scATAC-Seq, aberrant splicing
Single-Cell & CyTOF: Data gating, differential abundance
Statistics & AI
Statistical Modeling: Generalized Linear Models, Mixed Effects, MCMC, Time Series
Machine Learning: Neural Networks, Transformers, Random Forests, SVM, Clustering
Computer Vision & NLP: Clinical note extraction, segmentation, object detection