Deliverome Bio is a nonprofit startup building a first-of-its-kind open atlas of surface protein abundance and internalization. We’re hiring a founding team to generate the data that could expand what targeted therapies can reach.
The RoleYou will build and own the computational side of our proteomics platform: the pipelines, analysis, and data infrastructure that turn raw spectra from surface-enrichment experiments into quantitative measurements of the human surfaceome that we can compare across experiments and release publicly. This role is primarily scientist-facing. The people at the bench are who you serve first, and much of the work is building QC, reports, and result formats they can use directly without waiting on you. You will also contribute to our broader computational and data infrastructure, including how samples, metadata, and files are tracked, because the person who understands the data model should help decide how the data gets recorded in the first place.
We have an Evosep LC and an Orbitrap Astral Zoom, and we plan to run them at high sample throughput across many cell types and conditions. The data volumes and reprocessing demands that come with that need to be planned for from the beginning rather than handled once we are behind. We are flexible on whether compute lives in the cloud, on local hardware, or both. What we care about is scale, reproducibility a year later, and being able to release everything we generate. We currently run Spectronaut and are open to changing or adding to it, and we would value your perspective on the tradeoffs between commercial and open-source stacks at this throughput.
We also need someone who can diagnose the full workflow, not only the analysis. Surface proteomics can fail at labeling chemistry, enrichment specificity, digestion, chromatography, acquisition, search parameters, normalization, and statistics. We are looking for someone who can review a set of runs, say where the problem most likely sits, and propose the experiment that distinguishes between the remaining possibilities. That means knowing the wet lab well enough to push on experimental design while it is being decided, not after the samples are made.
This is a rare opportunity to build a public data resource from the first sample onward. No standard exists for how surfaceome abundance and internalization data should be quantified, QC'd, and compared across cell types and labs, and you will have unusual latitude to set one. Everything you build, including pipelines, spectral libraries, QC metrics, benchmarks, and the atlas itself, will be openly released and used by drug developers and academic labs worldwide. Like all of our founding hires, you will have a direct line to the co-founders and a genuine say in how the platform is built. You will gain practical, shareable expertise across large-scale MS analysis, research data infrastructure, and open data release that flexibly advances your career in either industry or academia.
What You Will DoAnalysis infrastructure at scale
Build, containerize, and maintain pipelines for bottom-up LC-MS/MS analysis, DIA and DDA, that run reproducibly across hundreds to thousands of files
Own the stack decisions: search and quantification software, workflow orchestration, compute and storage architecture, and cost
Build the data model and provenance layer, covering sample and run metadata, parameter capture, versioned reprocessing, and the ability to answer which pipeline version produced a given number a year later
Make reprocessing routine, so that improved methods are applied across the full corpus and not only to new data
Quantification, statistics, and the atlas
Design the quantification strategy for comparability across cell types and experiments, including normalization, batch structure, missing value handling, and protein inference and rollup
Implement differential and ranking statistics with error control that holds up in review and in reuse by others
Move toward absolute or ratio-anchored abundance estimates where the biology calls for it, and be explicit about the assumptions behind any copies-per-cell number
Explore data completeness across cell and tissue types to understand where more data is needed, or when aspects of the atlas are sufficiently statistically empowered
Build the atlas as a versioned, queryable resource with annotation and topology integration, confidence tiers, and a clear separation between what we measured and what we inferred
Scientist-facing tools and research infrastructure
Build QC reports, run-level dashboards, and result formats that let bench scientists check and interrogate their own data
Partner with lab operations on the infrastructure upstream of analysis, including sample and cell line tracking, electronic records, metadata capture at the point of experiment, and file handoff conventions
Own storage, backup, and data lifecycle for large instrument datasets, including what stays available, what gets archived, and what gets deposited
Contribute to our general computational tooling as needs come up, and be someone colleagues can bring a data problem to
Integration and open science
Deposit raw and processed proteomics data to ProteomeXchange/PRIDE or MassIVE with metadata that makes reuse practical, and facilitate deposition of sequencing data to GEO and SRA
Co-author methods and resource papers, and release pipelines and analysis code publicly
As we scale and publish our data, integrate proteomic abundance with pooled screening readouts, transcriptomic references, and public surfaceome resources to produce interpretable target rankings and open data viewers
Reanalyze published surfaceome and cell-surface capture datasets to benchmark our own results and to build on work the field has already done
What We Are Looking For
Required
PhD in computational biology, computer science, proteomics, or a related field, or equivalent depth built in industry
Deep hands-on experience analyzing large-scale bottom-up proteomics data, with fluency in at least one modern search and quantification stack (Spectronaut, DIA-NN, FragPipe/MSFragger, AlphaDIA, MaxQuant, Skyline). We run Spectronaut today, but experience matters more to us than the specific tool
Experience running analyses at scale, cloud or local, using containerized and orchestrated workflows (Nextflow, WDL, Snakemake, or similar)
Working command of quantitative proteomics statistics: normalization, batch effects, missing value structure in DIA versus DDA, protein inference, and FDR estimation
Proficiency in Python and version control with git/GitHub, including releasing tagged, reproducible pipelines
Highly valued
Experience with AWS (preferred) or equivalent, including object storage lifecycle management and batch or spot compute for large reprocessing jobs
Demonstrated ability to reason about the upstream experiment, including enrichment chemistry, sample preparation, LC, and instrument behavior, well enough to tell whether a failure is biological, chemical, or computational
Experience supporting bench scientists directly by building tools they can run themselves
Experience with electronic lab notebooks, LIMS, or sample tracking systems, and views on how to capture metadata in a busy lab
Building data resources that outside groups use: portals, APIs, versioned releases, FAIR metadata
Integration with orthogonal data types, including pooled screens, transcriptomics, membrane topology prediction, and existing surfaceome annotation sets
Experience with or interest in AI coding tools (Claude Code, Codex) as part of daily work
Who you are
You're a platform builder: you think about reproducibility, documentation, and what it takes for someone else to run your pipeline
You care about getting the measurement right, not just getting a number: you know what a given search setting, normalization choice, or FDR method assumes
You go upstream: when results look wrong you ask about the labeling reaction and the chromatography, not just the code
You build for the scientists around you, and would rather ship a tool a colleague can run than answer the same request every week
You know how to define the minimum viable analysis that answers this week's question while building toward a dataset that stays useful for a decade
You want your work to be adopted and reused by others, and your time to go towards problems that impact society
You thrive in a small, high-trust team where everyone's contributions are visible and decisions are made fast, and where accelerating science should never be perceived as stepping on toes
Level and Title
We are hiring at the Scientist through Principal Scientist level depending on capability and experience, and we prefer candidates who have built a platform or data product in industry. What matters is depth in large-scale MS data analysis, the judgment to make an analysis reproducible and releasable, and enough experimental fluency to tell us what isn't working and why.
What We OfferCompetitive salary benchmarked to industry rates
Full benefits including health, dental, and retirement
A founding team role: your decisions will shape the platform and the organization
Open science by design: your work will be published and used globally
Unprecedented collaboration with world-class advisors across proteomics, functional genomics, targeted delivery, and AI
A culture that values rigor, transparency, and speed, but doesn’t mistake busyness for productivity
Similar Jobs
What you need to know about the Boston Tech Scene
Key Facts About Boston Tech
- Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
- Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
- Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
- Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories



