Job description
As a Senior Data Engineer at Regard, you'll own the design, development, and production deployment of the data services that power the Regard platform. From ingesting and standardizing clinical data across health systems to making it reliably available for downstream product, analytics, and AI research workflows, you'll build and evolve the infrastructure that enables the platform. You'll help set the technical direction for the data team — designing lakehouse architecture, leading experimentation with processing and storage approaches, and mentoring engineers — while remaining hands-on with implementation and production delivery. We prioritize transparent, code-driven systems over black-box services, and you'll help architect the data platform that supports that philosophy.
About Regard
Regard's mission is to bring world-class healthcare to everyone. Our technology reasons through a patient's entire medical record to recommend diagnoses that would otherwise be missed, in real time at the point of care, during chart review, and in population-wide screening to identify patients who qualify for lifesaving treatment.
We work alongside some of the top health systems in the country to lead the change this industry needs. We're excited by challenges, mission-oriented work, and meaningful relationships. We want you to join us.
Our Tech Stack:
Lakehouse & processing: Python, SQL, Amazon S3 and S3 Tables, Apache Iceberg, PySpark, EMR Serverless, AWS Glue, Athena
Orchestration & infrastructure: Dagster, Kubernetes, Pulumi, GitLab CI/CD
Data services & analytics: PostgreSQL, ClickHouse, FastAPI, Metabase
Responsibilities:
Design and evolve the lakehouse architecture and shared processing capabilities that let the team build with an expanding range of data sources and use cases
Connect disparate clinical and product datasets into cohesive, reusable models that deepen understanding of clinician workflows and unlock new product and research opportunities
Lead continuous experimentation with Spark workloads, Iceberg storage, and orchestration — testing optimizations and measuring their effects on performance, scalability, and cost
Partner with Product and Clinical teams to identify valuable questions, define meaningful measures of adoption and impact, and deliver insights through visualization
Build data APIs, developer tooling, and self-service capabilities that make it easier for colleagues and AI assistants to discover data and create new applications
Partner with Research and Engineering on clinical AI workflows, including dataset curation, de-identification, human review, and reproducible model evaluation
Own projects from technical design through deployment and production operation, setting standards for data quality, documentation, and privacy
Guide engineers through design and code review, develop reusable platform capabilities, and create space for the team to experiment and learn
Minimum Qualifications:
Bachelor's degree in Computer Science, Mathematics, Statistics, a related field, or equivalent practical experience
5+ years of experience in data engineering roles, with significant hands-on experience using PySpark and cloud-based data services (AWS preferred)
Strong proficiency in Python and SQL, with a track record of delivering maintainable software through testing, code review, and CI/CD
Depth in data modeling and pipeline design, including integrating multiple sources and designing systems for scale and evolving use cases
Experience owning production systems and participating in on-call operational support
Track record of turning open-ended product or research questions into technical designs and communicating decisions and tradeoffs across technical and nontechnical teams
Preferred Qualifications:
Experience designing lakehouse systems with Apache Iceberg or similar table formats and developing reusable approaches to distributed data processing
Experience with Dagster, Kubernetes, infrastructure as code such as Pulumi, and AWS data catalogs and access controls such as Glue, Lake Formation, and IAM
Experience with healthcare data, HIPAA requirements, de-identification, and standards or vocabularies such as FHIR, OMOP CDM, or ICD-10
Experience partnering with Product to develop shared metrics, analytical datasets, and visualizations, or building data services with tools such as FastAPI, PostgreSQL, ClickHouse, or Athena
Practical experience with LLM-assisted development, including reviewing generated code and independently validating results
Experience supporting NLP or LLM systems through curated datasets, annotation and review workflows, or evaluation; familiarity with Model Context Protocol (MCP) integrations is a plus
Hybrid Work | Location | Work Authorization
For this role, Regard is currently only considering candidates who are authorized to work in the US without visa sponsorship, and are within the New York City, Los Angeles, or San Francisco metro areas
We expect our Engineers to be in the office on Tuesdays and Thursdays. We also require more frequent in-office work during the onboarding period and team onsite weeks up to once per month
We will provide relocation assistance to anyone who does not already reside in the NYC metro area
We prefer hiring people within commuting distance of our offices because we value getting together in person regularly
For those who enjoy working from our LA or Manhattan offices on a more regular basis, we offer catered lunches and other fun perks
Additionally, hybrid employees have the flexibility to work from locations outside of their home office from up to 6 weeks per year
Comp | Perks | Benefits
Eligible for equity
99% employer paid health benefits (Medical, Dental, and Vision) + One Medical subscription
18 PTO days/yr + 1 week holiday break
Monthly health & wellness budget
Company-sponsored team retreat + social events
A sabbatical program
Our goal at Regard is to provide and maintain a work environment that fosters mutual respect, professionalism and cooperation. Regard is proud to be an equal opportunity employer that does not discriminate on the basis of actual or perceived race, creed, color, religion, national origin, ancestry, alienage or citizenship status, age, disability or handicap, sex, gender identity, marital status, familial status, veteran status, sexual orientation or any other characteristic protected by applicable federal, state or local laws. We celebrate diversity and are proud of our supportive, inclusive workplace.
All candidates must successfully complete a background check as part of the hiring process.
Recruitment Fraud Notice
Regard only conducts hiring through official @regard.com email addresses. We will never ask candidates to pay fees, purchase equipment, or share sensitive personal information (such as Social Security numbers or bank details) during the interview process. If you receive an email about a job opportunity at Regard from a non-regard.com email address, it is not from us.