Skip to content
Our work / Data engineering 2023

Data engineering for a global apparel retailer: leaner Airflow pipelines and faster analytics on GCP

Overview

Dexis supported data engineering for a leading American apparel and lifestyle retailer with a large e-commerce footprint and operations across dozens of countries. The focus was the client’s Google Cloud estate - especially BigQuery and Apache Airflow - to cut redundant code, speed up processing, and make pipelines easier to extend.

Illustrative summary; client names and figures anonymised under NDA.

Data engineering case - panel 1: context and goals

Business context

The organisation opened BigQuery, an Airflow - based orchestration layer, development environments, and existing pipeline schemas to our engineers. The primary goal was to reduce the volume of code while improving performance and maintainability of batch and hybrid workloads feeding analytics and merchandising use cases.

Challenge

Legacy Airflow contained duplicated logic that slowed runs and made changes risky. The stack was hard to scale for new sources and domains, and dense, inconsistent code hurt readability - so onboarding and safe refactors were slower than the business needed.

Data engineering case - panel 2: challenges and tasks

Scope of work

Together with the client’s data team we prioritised: code optimisation (refactor for reuse and performance); data integration patterns around BigQuery loads; processing workflow design for streaming, batch, and parallel paths; algorithm and SQL review; secure data transfer and storage; performance tuning; and testing and validation in the target environment and Airflow - including automated checks with pytest for critical paths.

Used tools - BigQuery for fast analysis of large data volumes, Apache Zeppelin for interactive execution and reporting, Google Cloud Platform for compute and storage, Apache Airflow for workflow orchestration, Apache Spark for parallel processing, SQL for transformation, and pytest for automated testing of pipelines and Airflow tasks

Engineering & optimisation

We applied practical Spark and SQL patterns aligned with the deck’s playbook - for example favouring Apache Parquet where it reduced I/O, maximising parallelism in Spark, controlling shuffle, using broadcast hash joins where appropriate, caching intermediate results, and tuning executor memory. Apache Zeppelin supported collaborative exploration while we tightened queries and pipeline steps.

That combination moved the system from duplicated, brittle tasks toward templates and clearer SQL, improving both latency and operational clarity for the teams maintaining the estate.

Platform & tooling

The engagement centred on Google Cloud Platform services for compute and storage, BigQuery for large-scale analytics, Airflow for orchestration, Spark for parallel processing, Zeppelin for notebooks and visuals, and SQL as the main interface for transformation and optimisation - with pytest backing regression-style checks on pipelines and tasks.

Outcome

Engineers optimised a substantial body of code - on the order of twenty-plus developer environments’ worth of workload - validated using client cloud capacity to measure gains. The client expanded a large set of Zeppelin notebooks and highlighted strong collaboration, clear communication, and on-time, professional delivery.

Outcome of implementing data engineering - data integration into a single warehouse, data transformation and cleaning, data modeling for forecasting and segmentation, and business intelligence reporting through interactive dashboards Data pipelines automated to guarantee timely processing from source to destination, performance optimisation through real-time processing enabling personalised recommendations, and data governance and security policies for privacy, quality and CCPA and GDPR compliance Results - data engineers optimised over twenty working laptops with code hosted in the cloud, using the client's computing resources to measure the optimisation achieved; the client praised the team's expertise, cooperation and on-time delivery

Analytics & product data

Beyond pipeline mechanics, the programme sat alongside priorities common in retail: data integration from web, supply chain, sales, and reviews into a warehouse or lake; cleansing and modelling for forecasting and segmentation; BI and reporting for stakeholders; governed, compliant handling of customer data; and pathways toward near-real-time use cases such as recommendations - supported by a more dependable core ingestion and transformation layer.