Qualitate is building AI-powered moderation and structured customer insights, and this role establishes the data foundation behind those products. As Qualitate’s first dedicated Data Engineer, you will design the pipelines, models, and enrichment workflows that turn raw inputs into reliable, searchable outputs. You’ll also drive end-to-end data quality so enriched and extracted information is safe to use in production.
What you’ll do
- Build and maintain data pipelines: design ETL/ELT processes to ingest unstructured data, quantitative metrics, and 3rd party datasets into core systems.
- Design data models and systems: build and evolve data models across relational and non-relational environments.
- Create AI-powered enrichment pipelines: use LLMs and other AI systems to extract structured information from unstructured content, classify and enrich records, and integrate AI outputs into production workflows and evaluation systems.
- Own data quality end-to-end: implement automated QA checks, monitoring, and alerting to detect anomalies before they reach customers.
- Collaborate cross-functionally: partner with Research, Applied AI, and Engineering to define schemas, data contracts, and reporting needs as new features ship.
- Document as you build: maintain data lineage, schema documentation, and runbooks to support the systems you’re creating.
What you bring
- Strong SQL skills: ability to write complex, performant queries and reason about data at scale.
- Production engineering capability: experience with TypeScript or Python to build production-quality code for automation, pipelines, integrations, and internal tooling.
- Robust data integrations: working experience with APIs, databases, files, and external sources, including authentication, pagination, rate limits, incremental extraction, retries, schema changes, and error handling.
- Third-party dataset research and integration: experience evaluating company identifiers, entity-matching strategies, corporate hierarchies, subsidiaries, acquisitions, and mapping those entities across sources.
- Comfort across data types: familiarity with structured, semi-structured, and unstructured data, plus understanding tradeoffs between relational databases, object storage, search/indexing, and analytical stores.
- Alignment with current stack: support TypeScript pipelines running scheduled jobs and Postgres-native queues, S3 object storage, and AWS Bedrock for retrieval.
- Warehouse and transformation experience is a plus: e.g., Snowflake/BigQuery, dbt, and Airflow.
- Comfort with ambiguity: ability to build order in messy early-stage data infrastructure.
Technologies
SQL, TypeScript, Python, S3, AWS Bedrock, Postgres, Snowflake, BigQuery, dbt, Airflow
Compensation and benefits
- Base salary: USD 170,000 - 200,000 per year, depending on experience.
- Bonus: eligible for a cash bonus tied to individual and company performance.
- Equity: meaningful early-stage grant with upside tied to company growth.
- Benefits: health coverage, flexible PTO, and paid holidays.
- Location and schedule: full-time, in-person in New York City (onsite).
Additional fit signals
- Prior data engineering experience at a technology startup, ideally in fintech.
- Experience being early on the engineering team and helping build a data function from scratch.
- Exposure to LLM-driven data pipelines such as labeling, moderation, and quality scoring.
- Work history in market research, expert networks, or financial services, with an understanding of high data integrity expectations.
Qualitate emphasizes excellence, high velocity, and meaningful ownership. You’ll join a small early team focused on disrupting the multi-billion dollar market research industry with AI, with a collaborative approach to building systems that hold a high standard for quality and impact.