Shane Poole · Software & data engineer

Data is potential. Useful data is progress.

Five years of making data usable, then making it matter.

0yrs
Building production systems
0M
Documents processed
0+
Projects shipped
0
Papers co-authored

About

Nobody wants a pipeline. They want the answer at the end of it.

I'm a bioinformatics engineer on the data and development team at the Bove Lab in UCSF's Department of Neurology(opens in a new tab). The lab studies multiple sclerosis at the intersection of digital and precision medicine and sex and gender in neurology. Its research combines data from wearables, phones, and remote assessments with questions about how hormones and reproductive exposures shape inflammation and repair. I build the systems underneath that work, turning device streams, clinical records, and research data into tools and datasets the team can trust.

I tend to follow data all the way from source to decision. That might begin with an instrument in a clinic, a wearable, a phone, or a database of records. The destination might be an analysis or the dashboard a neurologist opens during a visit. That dashboard includes a map of local MS resources I helped build and years of gait and assessment videos arranged so a clinician can see how a patient's movement has changed. When the person represented by the data may be sitting in the room, quality stops being abstract. The same is true of extracting a structured measure from thousands of narrative notes: it can turn a question that was too expensive to ask into one a researcher can answer this quarter.

The assignment changes constantly. One week it is a cohort nobody has assembled before; the next, a figure for a grant due Friday. The underlying problems are familiar to any data team: sources that disagree, volumes that break the obvious approach, and results that still need to reproduce a year later. Owning systems over the long term has also taught me that the handoff matters as much as the build. I write the glossary, the runbook, and the decision record because a system nobody else can operate isn't finished.

I took an indirect route into software. Before UCSF, I spent three years as a chemist at the EPA's freshwater toxicology laboratory in Duluth, Minnesota, now home to the Great Lakes Toxicology and Ecology Division(opens in a new tab). The work moved between analytical chemistry, cell biology, animal studies, and the field. I prepared and measured chemical stocks for fish-tank exposures, ran LC-MS analyses, cell assays, and radioimmunoassays, and deployed fish in natural waters before retrieving, dissecting, and processing them. It gave me a literal source-to-result view of research: every measurement began with a chemical I prepared, an instrument I ran, or a fish I handled.

I wrote very little code there, but one VBA macro replaced an afternoon of copying and pasting in Excel with a keystroke. That feeling stuck. I went on to study software engineering at 42 Silicon Valley and never returned to the bench. Across both fields, I've co-authored 26 peer-reviewed papers, and my work has been cited over 450 times(opens in a new tab).

Languages

  • Python
  • SQL (T-SQL)
  • R
  • Dart / Flutter
  • Bash

Data

  • ETL & sharded pipelines
  • SQL Server, PostgreSQL
  • DuckDB + Parquet
  • Schema design
  • Statistical modeling

Services

  • FastAPI, Flask
  • Redis + RQ workers
  • Docker & Compose
  • REST API design
  • JWT auth

Platform

  • AWS (S3, RDS, ECR)
  • Object storage
  • nginx
  • CI & deployment
  • pytest

Selected work

012021 to present
Author & maintainer

A shared platform for research data

A Python toolkit that grew into our common data layer. It puts database access, LLM extraction from clinical notes, and geographic and socioeconomic enrichment behind one importable API, replacing one-off warehouse scripts with a consistent interface. Moving the read path to a local columnar store cut typical queries from more than 20 seconds to less than one. Five years and 30+ downstream notebooks later, it is the first import in nearly every analysis the group runs.

PythonDuckDB + ParquetSQL ServerLLM extraction
022026
Author & maintainer

Clinical text classification at scale

A multi-stage pipeline that classifies patient records by combining structured signals such as demographics, surgical history, diagnostic codes, and medications with rules-based NLP over free-text notes. A temporal ordering pass resolves the ambiguous records. Parallel shards, checkpoints, and resumable jobs keep processing practical at a scale of roughly 43 million documents.

PythonNLPSharded processingParquet
032021 to present
Primary developer

Multi-modal data capture platform

A platform that carries mobility research data from clinic hardware to an analysis-ready dataset. Instruments and a clinic-facing recording app send video and structured metrics to an asynchronous ingestion service, which validates and transcodes the files, records metadata in the study database, and moves media into object storage. A scheduled ETL process then reconciles identifiers across sources, giving downstream analysis one canonical layout instead of four conventions and a naming argument.

FlaskRedis + RQDockerObject storage
042023 to present
Pipeline author

Video analysis at half a terabyte

A pipeline that turns nearly 500 GB of gait and assessment video into measures researchers can analyze. It separates and transcribes the audio, then uses an LLM to extract structured measures from the transcripts. Computer vision tracks body pose and facial landmarks frame by frame. The facial analysis extracts action units and examines their combinations for patterns of emotional expression, while pose-derived gait measures are correlated with instrumented gait-mat outputs and other gold-standard clinical measures. The point is not simply to calculate features, but to validate them against established measures and make them useful in downstream research.

PythonComputer visionSpeech to textLLM extraction
052026
Backend lead

Participant video app and API

A mobile app and backend that let research participants record and submit standardized activity videos from their own phones. Designed for a compliance-restricted cloud environment, it sends resumable, chunked uploads through the API rather than exposing public storage. Authentication uses JWT and Argon2 password hashing, while interchangeable storage and mail adapters let the entire stack run locally with no cloud calls.

FastAPIFlutterPostgreSQLAWS
06Personal
Live

31: online card game

Live(opens in a new tab)

A real-time multiplayer version of the card game 31, built from scratch because my friends and I wanted to play online. I own the game logic, live-play experience, deployment, and infrastructure. It is a side project with real users, real uptime, and a real AWS bill.

Full-stackReal-timeAWSRoute 53

Also

LLM extraction pipelines
Built pipelines that extract structured clinical measures from narrative notes through an institutional Azure OpenAI gateway, with triplicate inference and inter-rater agreement scoring to test reliability.
Multi-site imaging infrastructure
Reconciled imaging manifests across seven sources and built a deterministic, de-identified ID service for a multi-site study.
Wearables and device data
Built sync services for consumer health APIs, VR cognitive assessment metrics, and gait instrumentation.
Analysis on request
Delivered 50+ cohort builds, statistical models, and publication figures through an ongoing internal service for collaborators across several labs.
Environmental modeling (EPA)
Modeled endocrine-disruption dose responses, worked with fish transcriptomic datasets, and contributed to adverse outcome pathway networks for assessing chemical-mixture risk in real watersheds.

Photography

Away from the keyboard.

Landscapes, night skies, and the occasional patient bird, shot on a Nikon D610. Mostly a hobby, though a few shoots for friends ended up on their wedding invitations. Click any frame to open it larger.

Contact

Let's talk.

If you're working on a difficult data problem, building tools for research, or simply want to compare notes, I'd be glad to hear from you. Email reaches me fastest.