Senior Databricks Engineer
USA
Job description
Position Description:
We are seeking an experienced Senior Databricks Engineer to serve as a hands on senior data engineer supporting a greenfield Databricks implementation. The role will design, develop, test, and deploy scalable batch and real time data engineering solutions while reverse engineering existing AWS based data processing to identify transformations, business rules, dependencies, and integration requirements that must be preserved or redesigned within the target Databricks architecture.
The Senior Databricks Engineer will deliver prioritized capabilities while helping establish reusable engineering patterns, development standards, governance, deployment practices, and best practices that improve the scalability, reliability, performance, and maintainability of the Databricks platform
This position can be located remotely anywhere in the U.S. Infrequent travel to Dallas, Texas may be required.
Key Responsibilities
. Design, develop, test, and deploy scalable Databricks data pipelines and transformation workflows as part of a greenfield Databricks implementation.
. Analyze and reverse engineer existing AWS based data processing solutions, including undocumented code and pipeline logic, to identify data transformations, business rules, dependencies, orchestration, and integration requirements that must be preserved or redesigned within Databricks.
. Build, enhance, and maintain Bronze, Silver, and Gold data layers supporting customer data models and downstream consumption requirements.
. Design and implement scalable ingestion solutions using Databricks Auto Loader, including schema management, checkpointing, incremental processing, backfill/reprocessing, and production scale file ingestion patterns.
. Develop and operate Lakeflow Declarative Pipelines (formerly Delta Live Tables) supporting batch and streaming workloads, data quality expectations, quarantine/error handling, pipeline dependencies, monitoring, and recovery.
. Develop real time and batch data processing solutions using Databricks, Structured Streaming, and associated technologies.
. Implement robust transformation logic using Apache Spark, PySpark, Spark SQL, and Delta Lake.
. Design and optimize Delta Lake solutions using appropriate MERGE patterns, Change Data Feed, partitioning/clustering, retention, and optimization strategies.
. Design solutions that support reliable ingestion, transformation, enrichment, and delivery of high volume customer data.
. Design and implement appropriate Unity Catalog governance structures, including catalogs, schemas, volumes, external locations, storage credentials, access controls, lineage, and promotion across development, test, and production environments.
. Develop and support integrations between Databricks, Adobe, AWS services, and other identified downstream systems.
. Analyze existing AWS Glue, Amazon S3, Amazon Redshift, Redshift Spectrum, and supporting AWS services to determine existing processing behavior, embedded business rules, dependencies, security requirements, and target state Databricks implementation.
. Establish reusable Databricks Workflows, deployment as code patterns, frameworks, utilities, and engineering standards appropriate for a greenfield implementation and subsequent development teams.
. Apply Databricks engineering best practices for data quality, performance optimization, scalability, observability, reliability, security, and maintainability.
. Perform performance tuning and optimization of Spark workloads, pipelines, queries, streaming processes, and data structures, including analysis of partitioning, shuffle behavior, join strategies, data skew, and Spark execution characteristics.
. Troubleshoot complex data pipeline, integration, data quality, performance, and production issues and participate in root cause analysis and remediation.
. Develop automated validation and testing approaches to ensure the accuracy, completeness, and reliability of data products.
. Perform legacy to target parity validation, reconcile discrepancies between AWS source processing and rebuilt Databricks pipelines, and confirm that required business rules and data outcomes have been preserved
. Evaluate legacy processing logic to distinguish required business functionality from source platform workarounds, avoiding unnecessary replication of legacy technical constraints within the target architecture.
. Participate in design discussions, peer code reviews, technical reviews, and solution refinement.
. Work within client established development, CI/CD, security, governance, and deployment processes.
. Produce and maintain technical documentation, including data flows, implementation details, development patterns, operational procedures, reverse engineering findings, and troubleshooting guidance.
. Provide hands on knowledge transfer and mentoring to client engineers to strengthen internal Databricks engineering capabilities and enable long term platform ownership.
. Leverage approved AI assisted development tools, where appropriate, to improve development velocity, code quality, testing, documentation, and engineering efficiency.
. Identify and recommend opportunities for automation and engineering improvement that increase delivery speed, solution quality, operational reliability, and recovery from production issues
Required Qualifications:
. 6+ years of data engineering or software engineering experience, including significant experience designing, developing, and supporting enterprise scale data platforms.
. 3+ years of hands on Databricks experience developing and operating production data engineering solutions.
. Demonstrated production experience with Databricks Auto Loader, including incremental cloud object storage ingestion, schema inference and evolution, schema hints, rescued data handling, checkpoint/state management, backfill and reprocessing strategies, and scalable file discovery approaches.
. Demonstrated production experience with Lakeflow Declarative Pipelines (formerly Delta Live Tables), including batch and streaming pipelines, data quality expectations, failed record/quarantine handling, streaming tables and materialized views, incremental versus full refresh processing, dependency management, monitoring, and alerting.
. Strong hands on experience implementing Unity Catalog in a production environment, including catalog/schema/volume design, external locations, storage credentials, grants, row and column level access controls, lineage, and asset promotion across environments.
. Strong hands on experience with Delta Lake, including MERGE patterns, Change Data Feed, time travel, OPTIMIZE, Z ORDER and/or liquid clustering, VACUUM and retention policies, and partitioning strategies at scale
. Strong hands on experience with Apache Spark, PySpark, and Spark SQL, including development of complex data transformations and production performance tuning.
. Demonstrated ability to diagnose and optimize Spark workloads using techniques including partition sizing, shuffle optimization, join strategy, skew handling, and Spark UI analysis.
. Experience developing production Structured Streaming solutions, including watermarking, late arriving data, stateful processing, checkpointing, and exactly once processing considerations.
. Demonstrated experience designing and implementing Medallion Architecture, including personally building Bronze, Silver, and Gold data layers and making appropriate decisions regarding responsibilities and boundaries between layers.
. Experience with Databricks Workflows and deployment as code, including job dependencies, retry/failure handling, alerting, and Git based environment promotion using Databricks Asset Bundles, Terraform, or comparable automation.
. Strong understanding of distributed data processing, data optimization, scalability, and production pipeline resiliency.
. Experience implementing production grade approaches for data quality, error handling, monitoring, logging, observability, and pipeline recovery.
. Experience with CI/CD, automated deployment practices, Git based source control, branching, pull requests, and peer code reviews.
AWS Source Environment Experience
The existing data processing environment is AWS based. Candidates must have sufficient hands on AWS data engineering experience to analyze existing implementations, understand data transformations and dependencies, and translate required functionality into the target Databricks architecture.
. Strong hands on experience developing and debugging AWS Glue ETL jobs using PySpark and Python, including DynamicFrames versus DataFrames, job bookmarks, connections, job parameters, worker sizing, and Glue runtime/version considerations.
. Demonstrated ability to analyze unfamiliar and undocumented AWS Glue code and determine the transformations, data movement, dependencies, and business rules being performed.
. Strong working knowledge of Amazon S3, including bucket/prefix structures, partitioning, file formats and compression, access patterns, and considerations affecting downstream data ingestion.
. Experience with Amazon Redshift, including data structures, distribution and sort strategies, COPY/UNLOAD patterns, stored procedures, and the ability to identify transformation or business logic embedded within the warehouse.
. Strong understanding of AWS IAM and data access patterns, including roles, assumed role relationships, policies, and authentication mechanisms sufficient to trace how existing data pipelines access source and target data and translate those requirements into appropriate Databricks and Unity Catalog access patterns
. Working knowledge of AWS Glue Crawlers and the Glue Data Catalog, including schema inference, schema drift, partition management, and the relationship between existing catalog definitions and the target Unity Catalog model.
. Working knowledge of Redshift Spectrum and external S3 backed data access patterns, with the ability to understand how existing external schemas and datasets should be represented within Databricks and Unity Catalog.
. Working knowledge of supporting AWS data, orchestration, monitoring, and security services such as Lambda, Step Functions, EventBridge, Athena, CloudWatch, and Secrets Manager. Candidates should be able to trace how these services participate in pipeline triggering, orchestration, monitoring, and credential management but are not expected to have expert level implementation experience with every service.
. Demonstrated ability to trace an end to end AWS data pipeline across multiple services, identifying data sources, transformations, orchestration, dependencies, security/access requirements, failure handling, and downstream consumers necessary to support migration to Databricks.
Reverse Engineering and Migration Experience
. Demonstrated experience reverse engineering undocumented legacy data pipelines where business rules and processing requirements are embedded within application, ETL, orchestration, or database code.
. Experience identifying pipeline behavior and dependencies where original developers or complete technical documentation are unavailable.
. Demonstrated experience performing data reconciliation and parity validation between legacy and rebuilt pipelines, investigating discrepancies and confirming equivalent business outcomes.
. Ability to distinguish business logic that must be preserved from technical workarounds or constraints of the legacy platform, and redesign appropriately for a modern Databricks architecture.
. Demonstrated ability to troubleshoot complex data engineering and production issues independently while collaborating within an integrated delivery team.
. Experience working in Agile delivery environments and delivering against prioritized product backlogs.
. Strong communication and collaboration skills with the ability to work effectively across engineering, architecture, product, and business teams.
Key Skills
Databricks | Auto Loader | Lakeflow Declarative Pipelines | Unity Catalog | Delta Lake | Apache Spark | PySpark | Spark SQL | Structured Streaming | Medallion Architecture | Bronze/Silver/Gold Data Layers | Databricks Workflows | Databricks Asset Bundles | Terraform | Real Time Data Processing | Data Pipelines | ETL/ELT | Data Quality | Performance Tuning | AWS Glue | Glue Data Catalog | Amazon S3 | Amazon Redshift | Redshift Spectrum | AWS IAM | Lambda | Step Functions | EventBridge | Athena | CloudWatch | CI/CD | Git | Adobe Integration | AWS Cloud Data Engineering | Data Migration | Reverse Engineering | Data Reconciliation | Agile Delivery
CGI is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. To support the ability to reward for merit based performance, CGI typically does not hire individuals at or near the top of the range for their role. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $100,800.00 $245,500.00.
CGI's benefits are offered to eligible professionals on their first day of employment to include:
. Competitive compensation including profit participation program
. Comprehensive medical, dental, and vision benefits
. Basic life and accidental death & dismemberment insurance
. Matching contributions through 401(k) plan, and CGI share purchase plan
. Flexibility and paid accrued vacation leave, ranging from 10 to 20 days per year, based on job level, years of relevant prior experience, and years of service
. 10 paid holidays per year
. At least 80 consecutive hours of paid sick/safe leave (except where applicable state/local law requires more)
. Paid parental leave, ranging from 20 to 70 consecutive business days based on circumstances of leave and applicable laws
. Bereavement leave, ranging from 1 to 7 days per year based on relationship.
. Paid jury duty leave, up to time summoned
. Learning opportunities and tuition assistance
. Wellness and Well being programs
For more detailed information about our benefits offerings visit Benefits | CGI Careers
Please note that the benefits listed above are subject to change based on the specific terms and conditions of the contract being supported.
CGI anticipates accepting applications for this position through 2026 09 18.
Skills:
· Agile
· Amazon Web Services Cloud
· Apache Spark
· Databricks
· ETL
· GIT
· Performance Tuning
· SQL
What you can expect from us:
Together, as owners, let’s turn meaningful insights into action.
Life at CGI is rooted in ownership, teamwork, respect and belonging. Here, you’ll reach your full potential because…
You are invited to be an owner from day 1 as we work together to bring our Dream to life. That’s why we call ourselves CGI Partners rather than employees. We benefit from our collective success and actively shape our company’s strategy and direction.
Your work creates value. You’ll develop innovative solutions and build relationships with teammates and clients while accessing global capabilities to scale your ideas, embrace new opportunities, and benefit from expansive industry and technology expertise.
You’ll shape your career by joining a company built to grow and last. You’ll be supported by leaders who care about your health and well-being and provide you with opportunities to deepen your skills and broaden your horizons.
Come join our team—one of the largest IT and business consulting services firms in the world.
Qualified applicants will receive consideration for employment without regard to their race, ethnicity, ancestry, color, sex, religion, creed, age, national origin, citizenship status, disability, pregnancy, medical condition, military and veteran status, marital status, sexual orientation or perceived sexual orientation, gender, gender identity, and gender expression, familial status or responsibilities, reproductive health decisions, political affiliation, genetic information, height, weight, or any other legally protected status or characteristics to the extent required by applicable federal, state, and/or local laws where we do business.
CGI provides reasonable accommodations to qualified individuals with disabilities. If you need an accommodation to apply for a job in the U.S., please email the CGI U.S. Employment Compliance mailbox at US_Employment_Compliance@cgi.com . You will need to reference the Position ID of the position in which you are interested. Your message will be routed to the appropriate recruiter who will assist you. Please note, this email address is only to be used for those individuals who need an accommodation to apply for a job. Emails for any other reason or those that do not include a Position ID will not be returned.
We make it easy to translate military experience and skills! Click here to be directed to our site that is dedicated to veterans and transitioning service members.
All CGI offers of employment in the U.S. are contingent upon the ability to successfully complete a background investigation. Background investigation components can vary dependent upon specific assignment and/or level of US government security clearance held. Dependent upon role and/or federal government security clearance requirements, and in accordance with applicable laws, some background investigations may include a credit check. CGI will consider for employment qualified applicants with arrests and conviction records in accordance with all local regulations and ordinances.
CGI will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by the employer, or (c) consistent with CGI’s legal duty to furnish information.