Job description
Job Family : Software Development & Support Travel Required : None Clearance Required : Ability to Obtain Public Trust AWS Lakehouse Data Engineer We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments.
This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.
What You Will Do
- Build and Operate Data Pipelines (Batch and Streaming) Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
- Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.
- Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
- Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
- Deliver an AWS-Native Lakehouse Data Platform Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
- Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
- Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
- Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
- Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.
- Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.
- Metadata, Governance, Access Control, Lineage, and Quality Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services.
- Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.
- Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.
- Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling.
- Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.
- AWS Automation, CI/CD, and Operations Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.
- Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.
- Implement observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.
- Continuously evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.
- Cross-Team Collaboration and Documentation Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines.
- Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines.
- Present technical findings, trade-offs, risks, and recommendations clearly to technical and non-technical stakeholders.
- What You Will Need Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in leu of degree.
- SIX (6) years of relevant experience.
- Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
- Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
- Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
- Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
- Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
- Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
- Experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
- Experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
- Ability to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.
- What Would Be Nice to Have Hands-on experience with Databricks, Delta Lake, or migrating Databricks workloads to AWS-native services and Apache Iceberg.
- Experience with AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), Amazon Kinesis, AWS Database Migration Service (DMS), AWS Lambda, Amazon MSK, or similar ingestion and orchestration services.
- Experience with modern DevOps practices and tools such as Git, Terraform, AWS CloudFormation or AWS CDK, Jenkins, AWS CodePipeline, GitHub Actions, and Docker.
- Experience integrating lakehouse data with business intelligence and visualization tools such as Amazon QuickSight, Tableau, or Power BI.
- Experience using AI-assisted coding tools, such as GitHub Copilot, ChatGPT, Cursor, or Kiro, to accelerate implementation while maintaining code quality, testing, review, privacy, and security controls.
- Knowledge graph and Graph RAG experience, including graph modeling, ontology and taxonomy alignment, entity resolution, relationship extraction, and hybrid retrieval that combines graph traversal with semantic or vector search.
- The annual salary range for this position is $113,000.00-$188,000.00.
Compensation
decisions depend on a wide range of factors, including but not limited to skill sets, experience and training, security clearances, licensure and certifications, and other business and organizational needs.
What We Offer
Guidehouse offers a comprehensive, total rewards package that includes competitive compensation and a flexible benefits package that reflects our commitment to creating a diverse and supportive workplace.
Benefits
- include: Medical, Rx, Dental & Vision Insurance Personal and Family Sick Time & Company Paid Holidays Parental Leave 401(k) Retirement Plan Group Term Life and Travel Assistance Voluntary Life and AD&D Insurance Health Savings Account, Health Care & Dependent Care Flexible Spending Accounts Transit and Parking Commuter Benefits Short-Term & Long-Term Disability Tuition Reimbursement, Personal Development, Certifications & Learning Opportunities Employee Referral Program Corporate Sponsored Events & Community Outreach Care.com annual membership Employee Assistance Program Supplemental Benefits via Corestream (Critical Care, Hospital Indemnity, Accident Insurance, Legal Assistance and ID theft protection, etc.)
- Position may be eligible for a discretionary variable incentive bonus About Guidehouse Guidehouse is an Equal Opportunity Employer–Protected Veterans, Individuals with Disabilities or any other basis protected by law, ordinance, or regulation.
- Guidehouse will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of applicable law or ordinance including the Fair Chance Ordinance of Los Angeles and San Francisco.
- If you have visited our website for information about employment opportunities, or to apply for a position, and you require an accommodation, please contact Guidehouse Recruiting at 1-571-633-1711 or via email at RecruitingAccommodation@guidehouse.com .
- All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodation.
- All communication regarding recruitment for a Guidehouse position will be sent from Guidehouse email domains including @guidehouse.com or guidehouse@myworkday.com .
- Correspondence received by an applicant from any other domain should be considered unauthorized and will not be honored by Guidehouse.
- Note that Guidehouse will never charge a fee or require a money transfer at any stage of the recruitment process and does not collect fees from educational institutions for participation in a recruitment event.
- Never provide your banking information to a third party purporting to need that information to proceed in the hiring process.
- If any person or organization demands money related to a job opportunity with Guidehouse, please report the matter to Guidehouse’s Ethics Hotline.
- If you want to check the validity of correspondence you have received, please contact recruiting@guidehouse.com .
- Guidehouse is not responsible for losses incurred (monetary or otherwise) from an applicant’s dealings with unauthorized third parties.
- Guidehouse does not accept unsolicited resumes through or from search firms or staffing agencies.
- All unsolicited resumes will be considered the property of Guidehouse and Guidehouse will not be obligated to pay a placement fee.
