Data Engineer (Databricks), Assistant Vice President
Job Description
State Street is seeking a Data Engineer (Databricks) to serve as an Assistant Vice President, focused on designing, building, and supporting a Legal Data Lakehouse platform on AWS and Databricks. The role emphasizes scalable data pipelines, governance, and trusted data capabilities for legal operations, compliance analytics, reporting, and AI/ML use cases.
Key Responsibilities
- Design, build, and maintain scalable data pipelines using PySpark, Python, and Spark SQL
- Develop and optimize ETL/ELT workflows on Databricks using Delta Lake
- Implement lakehouse architecture using Bronze/Silver/Gold layers for enterprise data platforms
- Build and manage Databricks Jobs, Workflows, and Notebooks for batch and streaming workloads
- Create reusable frameworks for data ingestion, processing, and orchestration
- Containerize data workloads with Docker and automate processes through scripting
- Integrate Databricks pipelines with Power Platform solutions including Power Apps and Power Automate
- Enable data exposure for business users via APIs, connectors, and curated datasets
- Integrate data from SQL Server, Oracle, and other enterprise source systems
- Apply advanced data modeling techniques, including dimensional modeling and data partitioning/optimization
- Work with structured and semi-structured data such as JSON and Parquet
- Tune performance using caching, indexing, and Spark optimization approaches
- Publish curated datasets for consumption in Power BI dashboards and Power Apps/Power Automate workflows
- Ensure data quality with unit testing, validation frameworks, and automated checks
- Monitor and troubleshoot distributed Spark workloads across Databricks and cloud environments; analyze logs to resolve production issues
- Maintain data lineage, consistency, and audit readiness
- Collaborate with Legal, Security, Compliance, and Enterprise Data teams to deliver scalable solutions
- Translate business requirements into robust data engineering designs
- Serve as a Subject Matter Expert (SME) in Databricks and lakehouse architecture
- Lead initiatives with minimal supervision and take full ownership of deliverables
- Support data governance using Databricks Unity Catalog and AWS controls such as IAM and KMS
- Ensure compliance with data privacy and regulatory requirements including GDPR, as well as internal security and audit standards
- Design and maintain data access controls and data classification and handling standards; collaborate with IAM and security teams for secure access
- Design and maintain CI/CD pipelines using Harness, Azure DevOps, or GitHub
- Automate deployment of Databricks assets using Databricks Repos and Databricks CLI
- Monitor, schedule, and optimize workflows using Databricks orchestration tools
- Maintain clear documentation including architecture, data flows, and runbooks
- Continuously improve performance, scalability, and cost efficiency
Required Qualifications
- 8+ years of experience in Data Engineering or data platform development
- Strong hands-on experience with Databricks and Apache Spark
- Proficiency in PySpark, Python, and SQL
- Experience with AWS data platform services including S3, Glue, Lambda, and IAM
- Experience with Delta Lake and lakehouse architecture
- Solid understanding of distributed data processing
- Solid understanding of ETL/ELT frameworks
- Solid understanding of data modeling techniques
Education
Bachelor's or Master's degree in computer science, Data Engineering, Information Systems, or a related technical discipline.
Technologies
Databricks, AWS, PySpark, Python, Spark SQL, ETL/ELT, Delta Lake, Lakehouse architecture (Bronze/Silver/Gold layers), Databricks Jobs, Databricks Workflows, Notebooks, Docker, Power Platform (Power Apps, Power Automate), APIs, Connectors, Power BI, SQL Server, Oracle, JSON, Parquet, Unit testing, Validation frameworks, Databricks Unity Catalog, IAM, KMS, GDPR, CI/CD pipelines, Harness, Azure DevOps, GitHub, Databricks Repos, Databricks CLI, S3, Glue, Lambda
Preferred Qualifications
- Hands-on experience with Databricks platform components including Delta Lake, Workflows, and Unity Catalog
- Strong experience building end-to-end data pipelines (batch and streaming) using AWS and Databricks
- Familiarity with performance optimization techniques in Spark and Delta Lake
- Experience supporting analytics, reporting, or AI/ML use cases on a lakehouse platform
- Understanding of data governance, metadata management, and security controls
Nice to Have
- Experience in Legal, Compliance, Financial Services, or regulated industries
- Understanding of legal data constructs such as contracts, clauses, obligations, and matters
- Exposure to unstructured data processing or document/NLP pipelines
- Experience with Power BI, Power Apps, or Power Platform
- Experience handling sensitive data in audit-driven environments
Benefits
- 401(k) with company match
- Insurance coverage including basic life, medical, dental, vision, long-term disability, and other optional additional coverages
- Paid-time off including vacation, sick leave, short term disability, and family care responsibilities
- Access to the Employee Assistance Program
- Incentive compensation including eligibility for annual performance-based awards (excluding certain sales roles subject to sales incentive plans)
- Eligibility for certain tax advantaged savings plans
Company Values
- Strong analytical thinking and problem-solving skills
- Hands-on expertise in data engineering and distributed data processing
- Ability to work in a fast-paced, enterprise environment
- Effective communication and collaboration with cross-functional teams
- Ownership mindset with focus on quality, scalability, and performance
Role Details
Location: Quincy, MA (onsite)
Compensation: USD 110,000 - 177,500 per yearly
Experience: Minimum 8 years