Decagon is seeking a Senior Data Infrastructure Engineer to design, build, and operate the data systems powering its AI products. This onsite role in San Francisco, CA will own end-to-end data pipelines and storage, driving reliability and performance and enabling engineers to work with data at scale.
Responsibilities
- Design and implement high-throughput data pipelines and streaming systems with strong SLOs, clear runbooks, and actionable telemetry.
- Build and operate real-time and batch ingestion infrastructure using tools like Kafka, Flink, and Airflow.
- Own our analytical data layer — schema design, query performance, and cost optimization across ClickHouse, BigQuery, or similar.
- Partner with research and product teams to architect data solutions, evaluate performance, and scale new features.
- Tune pipeline and query latencies: optimize data paths, apply smart caching/partitioning, and hit tight p95/p99 targets.
- Lead infrastructure-as-code (Terraform) and GitOps practices for data systems; reduce drift with reusable modules and policy-as-code.
- Participate in on-call and drive down toil through automation and elimination of recurring data issues.
Requirements
- 5+ years building and operating production data infrastructure at scale.
- Hands-on experience with Tier 1 data technologies: ClickHouse, Kafka (or MSK/Pub-Sub/RabbitMQ), and Flink or dbt.
- Proven track record meeting high availability and low latency targets across streaming and batch workloads.
- Excellent observability chops (OpenTelemetry, Prometheus/Grafana, Datadog) and strong incident response discipline.
- Clear written communication and the ability to turn ambiguous data requirements into simple, reliable designs.
Technologies
- Kafka, MSK, Pub/Sub, RabbitMQ
- Flink, Airflow, dbt
- ClickHouse, BigQuery
- Terraform, Kubernetes, GKE, EKS, AKS
- Snowflake, Redshift, Databricks
- Debezium, Dagster, Prefect
- Spark, Dask
- OpenTelemetry, Prometheus, Grafana, Datadog
- GCP, AWS, Azure
Benefits
- Take what you need vacation policy (subject to local requirements; UK employees receive 25 days of statutory leave)
- Medical, Dental, and Vision benefits for you and your family
- Life Insurance and Disability Benefits
- Retirement Plan (e.g., 401K, pension)
- Parental Leave
- Fertility and family building benefits through Carrot
- Daily lunches and snacks in the office to keep you at your best
About the Team
The Infrastructure team builds and operates the foundations that power Decagon, including networking, data, ML serving, the developer platform, and real-time voice. We partner closely with product, data, and ML to deliver high-scale, low-latency systems with clear SLOs and strong developer ergonomics.
- Focus areas include Core Infra, Data Infra, ML Infra, and Platform (DevEx).
- Core Infra covers the foundational cloud stack—networking, compute, storage, security, and infrastructure-as-code for reliability, scale, and cost efficiency.
- Data Infra powers analytics, BI, and customer-facing telemetry across environments.
- ML Infra provides GPU and model-serving platforms for LLM inference with multi-provider routing and on-prem capabilities.
- Platform (DevEx) focuses on CI/CD, paved paths, and core services that accelerate shipping with safety and consistency.
About the Role
You will design, build, and operate the data systems that power Decagon's AI products, owning critical data pipelines and storage end-to-end. The role emphasizes improving reliability and performance and creating paved paths that let every Decagon engineer work confidently with data at scale.
Compensation
Compensation: $200K – $400K + Equity. This range reflects the expected compensation for this role. Compensation within the range is determined based on experience, skills, and the scope of responsibilities, with flexibility for candidates who demonstrate exceptional impact.