Case Study
A Scalable Foundation for Growth
5 upstream systems, 4 source formats, and 3 data layers unified into one automated foundation—supporting daily SLA commitments, trusted analytics, stronger governance, and faster onboarding of future data sources.
Established in the mid-1800s, the enterprise has a long-standing commitment to helping people prepare for the future. Today, it continues to empower its customers to make the most of their financial future, building on its experience and commitment to supporting them through every stage of their financial journey.
As part of a highly regulated industry, the enterprise relies on accurate, timely, and governed data to support investment operations, business decision-making, regulatory reporting, and compliance requirements.
The customer operates at an enterprise scale, managing data from multiple upstream systems and serving a broad set of downstream analytics and reporting stakeholders. Investment-related data was received from five different upstream systems in heterogeneous formats including CSV, XLSX, XLSB, and TXT, creating a complex data environment that needed to support daily data delivery requirements and timely availability of curated datasets.
The customer needed a robust, scalable, and automated cloud-native data platform that could support future growth while maintaining cost efficiency.
The customer faced challenges in managing and processing investment-related data across five different upstream systems and heterogeneous formats including CSV, XLSX, XLSB, and TXT, making standardized processing difficult.
There was no standardized data architecture separating raw, processed, and business-ready datasets. Inconsistent source structures also resulted in data quality and usability challenges, making it difficult to establish clear traceability and increasing effort for downstream reporting and analytics.
Data pipelines relied heavily on manual intervention, leading to operational inefficiencies and increased risk of failures. Manual processing also increased engineering effort for data corrections and reprocessing activities.
Limited monitoring and observability made it difficult to identify and resolve pipeline issues proactively, increasing the risk of downtime and operational disruption across daily processing workflows. The fragmented processing environment also created higher operational risk in meeting daily SLA commitments, with delays potentially affecting downstream reporting and analytics processes.
Data quality and processing challenges reduced confidence in data used for business and regulatory decision-making. In a regulated financial environment, this created potential governance and compliance concerns, while increasing overall operational and regulatory risk.
Hexaware designed and implemented a cloud-native data engineering platform on Google Cloud Platform, specifically tailored to meet the customer’s regulatory, operational, and scalability requirements. This Google Cloud data engineering-led foundation provided the customer with a robust, scalable, and automated cloud-native data platform.
The customer’s investment data was organized through a medallion architecture (ODP → FDP → CDP), creating a clear separation of raw, cleansed, and business-ready datasets. This standardized approach improved data processing, quality, traceability, and data governance, while supporting downstream analytics.
Data from five upstream systems in CSV, XLSX, XLSB, and TXT formats was landed into Google Cloud Storage (GCS). Apache Airflow DAGs orchestrated end-to-end ingestion workflows, while Dataflow pipelines built using Apache Beam and Python processed structured and semi-structured source files.
Data was first loaded into staging tables and subsequently ingested into ODP, providing the customer with a consistent raw data layer before further cleansing, standardization, and integration.
The ODP layer provided a clear foundation for raw data, enabling source data to be retained before further processing.
In the FDP Layer (Silver), data was cleansed, standardized, and integrated. SQL-based transformation logic was applied on ODP datasets, with data quality and consistency checks incorporated during processing.
In the CDP Layer (Gold), business-ready datasets were created through further aggregation and enrichment. Data models were optimized for reporting, analytics, and downstream consumption.
This layered approach improved traceability and auditability by maintaining a clear separation between raw, cleansed, and business-ready datasets.
Apache Airflow managed dependencies across all processing stages. Modular DAG design enabled reusability, failure isolation, simplified maintenance, and future scalability.
This data pipeline automation reduced manual intervention and operational overhead while providing scalable workflow management with improved failure handling.
The customer’s business-ready data models were optimized for reporting, analytics, and downstream consumption. BigQuery performance and cost optimization techniques were incorporated to improve processing performance while controlling cloud costs.
Developing Improved Visibility and Operational Reliability
Log-based metrics were implemented within GCP, while automated email notifications were configured for failures and anomalies. The customer gained enhanced observability across data processing workflows, improving operational reliability and reducing downtime.
Terraform was used to provision and manage all cloud resources, with infrastructure maintained through version-controlled deployments. This enabled consistent deployments across development, testing, and production environments and supported consistency, repeatability, and data governance across environments.
The architecture was specifically tailored to meet the customer’s regulatory, operational, and scalability requirements. The layered medallion architecture aligned with financial services data governance standards, while clear separation of data layers improved traceability and auditability.
Automated end-to-end processing reduced manual intervention and operational overhead. Reusable pipeline components supported rapid onboarding of future data sources, providing greater scalability, operational resilience, data quality, and long-term maintainability compared to fragmented and manually managed pipelines.
A scalable, automated, and governed Google Cloud data platform improving data quality, operational reliability, reporting readiness, and confidence in enterprise data.
The platform supports daily SLA-driven data delivery requirements through automated end-to-end processing and improved workflow management.
Data from five different upstream systems and heterogeneous file formats can be processed through a standardized data architecture, improving consistency and simplifying future onboarding.
The ODP-FDP-CDP layered architecture provides clear separation of raw, cleansed, and business-ready datasets, improving traceability, auditability, and data governance.
Automated end-to-end processing reduced manual intervention and operational overhead, while modular Airflow workflows improved failure isolation and maintenance.
Data quality and consistency checks incorporated during processing support improved data completeness across source systems and reduce data processing errors and reprocessing activities.
Curated datasets are made available for reporting and analytics teams, improving the availability of trusted datasets and providing faster access to business-ready information.
Log-based metrics, automated failure notifications, and enhanced observability improved operational reliability, reduced downtime, and strengthened the ability to detect and resolve pipeline failures.
Terraform-based infrastructure management ensures consistency, repeatability, and governance across environments, while reusable pipeline components support rapid onboarding of future data sources. The Google Cloud platform provides a scalable foundation for future growth.
The enterprise is now better positioned to support future growth with a scalable, automated, and governed data foundation on Google Cloud.
With data pipeline automation, standardized ODP-FDP-CDP Medallion architecture, reusable pipeline components, and strong data governance, the enterprise can onboard new data sources faster, expand analytics capabilities, support daily SLA commitments, and build greater confidence in its enterprise data assets.
Its future expansion can support broader analytics capabilities and faster access to business-ready information. Reusable components provide a foundation for continued growth while maintaining governance and operational reliability. The platform is positioned to evolve with the enterprises’ changing investment data needs.
See how Hexaware’s Google Cloud partnership can help build a modern data foundation and prepare your organization for what’s next.