Home

Samara reddy - Data Engineer
[email protected]
Location: Farmington Hills, Michigan, USA
Relocation: YES
Visa: H1B
Resume file: Samara_Reddy (2)_1789481176947.docx
Please check the file(s) for viruses. Files are checked manually and then made available for download.
Samara Reddy
Email: [email protected]
PH: +1(248) 812-9311
Senior Data Engineer

Professional Summary:

Seasoned Senior Cloud Data Engineer with 7+ years of experience designing, developing, and operating enterprise data platforms across Azure and AWS, with exposure to GCP.
Advanced expertise in Snowflake data engineering, cloud data warehousing, ELT architecture, dimensional modeling, workload optimization, and enterprise migration programs.
Hands-on proficiency with dbt for modular SQL transformations, models, incremental processing, snapshots, macros, seeds, testing, documentation, and reusable ELT workflows.
Strong Azure background spanning Azure Data Factory (ADF), Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Synapse Analytics, Azure SQL Database, Microsoft Fabric, OneLake, and Azure Databricks.
Expertise in Databricks, Apache Spark, PySpark, Delta Lake, Delta Live Tables (DLT), Unity Catalog, Structured Streaming, Spark SQL, and lakehouse architecture.
Advanced SQL capability covering complex joins, CTEs, window functions, stored procedures, query optimization, execution-plan analysis, indexing, partitioning, and performance tuning.
Experienced in building scalable ETL/ELT pipelines with source-to-target mapping, data transformation, reconciliation, validation, incremental loading, error handling, orchestration, and production support.
Proven ability to create enterprise data models using Kimball methodology, star schema, snowflake schema, fact and dimension structures, conformed dimensions, and SCD Type 1/Type 2.
Experienced in Snowflake performance engineering including virtual warehouse sizing, clustering, micro-partitioning, caching, workload management, concurrency optimization, and credit-aware processing.
Hands-on knowledge of Snowflake objects and capabilities including databases, schemas, stages, tables, views, materialized views, secure access, monitoring, data sharing, and governed consumption.
Implemented cloud migration and modernization initiatives involving Teradata, Hadoop, Hive, Informatica, Oracle, SQL Server, legacy ETL platforms, Snowflake, Databricks, and cloud-native lakehouses.
Built batch and streaming data solutions using Azure Data Factory, Databricks Workflows, DLT, Spark Structured Streaming, Kafka, AWS Glue, Lambda, EMR, Kinesis, and event-driven processing patterns.
Experienced with data governance, metadata management, lineage, cataloging, data quality, RBAC, RLS, encryption, auditability, Microsoft Purview, Unity Catalog, and regulated-data controls.
Proficient in CI/CD and DataOps using Azure DevOps, GitHub Actions, Git, Terraform, Azure Bicep, automated testing, release approvals, infrastructure as code, and deployment automation.
Built governed lakehouse architectures using bronze, silver, and gold layers with Delta Lake and Apache Iceberg, supporting schema evolution, ACID transactions, time travel, and scalable analytics.
Experienced in developing Power BI semantic models, DAX measures, KPI dashboards, incremental refresh, aggregations, query diagnostics, and self-service analytics for business stakeholders.
Applied Python, PySpark, Scala, SQL, Shell, and Java for data processing, validation, automation, ingestion utilities, transformation logic, and distributed computing workloads.
Processed high-volume and petabyte-scale datasets using Spark, EMR, ADLS Gen2, S3, Hive, HDFS, and cloud warehouse technologies while improving reliability, latency, and operational efficiency.
Experienced with automated data-quality frameworks using Great Expectations, DLT expectations, Python/SQL validation, profiling, anomaly detection, reconciliation, and migration controls.
Implemented monitoring and observability using Azure Monitor, Log Analytics, Application Insights, Databricks monitoring, Snowflake monitoring, operational alerts, and root-cause analysis.
Healthcare data experience supporting HIPAA-aligned pipelines, EDI 837/835/270/271/276/277, claims and member data, secure processing, auditability, and controlled access.
Experienced in relational and NoSQL technologies including Snowflake, Azure SQL Database, SQL Server, Oracle, PostgreSQL, Teradata, DB2, MySQL, MongoDB, Cassandra, and HBase.
Worked with semi-structured and open data formats including JSON, Parquet, Avro, ORC, Delta, and Iceberg for ingestion, standardization, transformation, schema management, and analytics.
Effective Agile/Scrum contributor with experience in requirements analysis, technical design, stakeholder collaboration, documentation, code reviews, UAT, production deployment, incident resolution, and mentoring.

IT SKILLS:
Big Data: Cloudera Distribution, HDFS, YARN, MapReduce, Pig, Sqoop, Kafka, HBase, Hive, Flume, Cassandra, Spark, Spark Streaming, Storm, Scala, Impala
Programming: Python, PySpark, Scala, Java, C, C++, Shell Scripting, Perl, SQL, PL/SQL, Bash
Data Engineering: Snowflake, dbt, Databricks, Delta Lake, Delta Live Tables, Apache Iceberg, Microsoft Fabric, ETL/ELT, CDC, Data Integration, Data Quality, Data Validation
Snowflake: Architecture, Virtual Warehouses, Databases, Schemas, Stages, Tables, Views, Materialized Views, Micro-Partitioning, Clustering, Caching, Workload Management, Cost Optimization, Secure Data Sharing
Azure: Azure Data Factory, ADLS Gen2, Azure Synapse Analytics, Azure SQL Database, Azure Databricks, Microsoft Fabric, OneLake, Event Hubs, Azure Functions, Azure Monitor, Log Analytics, Purview, Azure DevOps
AWS: S3, EMR, Glue, Lambda, Kinesis, DMS, Redshift, Aurora, RDS, DynamoDB, EC2, Lake Formation, Glue Data Catalog, Athena, QuickSight, Boto3
dbt: Models, Incremental Models, Snapshots, Macros, Seeds, Tests, Documentation, Source-to-Target Transformations, Reusable ELT Workflows
ETL / Integration: IBM DataStage, Informatica PowerCenter, Ab Initio, Azure Data Factory, AWS Glue, Sqoop, Source-to-Target Mapping, Reconciliation, Migration Validation
Database Technologies: Snowflake, Azure SQL Database, SQL Server, Oracle, PostgreSQL, Teradata, IBM DB2, MySQL, NoSQL, MongoDB, Cassandra, HBase
Data Modeling: Kimball Methodology, Dimensional Modeling, ER Modeling, Star Schema, Snowflake Schema, Fact Tables, Dimension Tables, Conformed Dimensions, SCD Type 1/Type 2
SQL Engineering: Advanced SQL, Query Optimization, Execution Plans, CTEs, Window Functions, Stored Procedures, Functions, Triggers, Indexing, Partitioning, Performance Tuning
Data Governance: Microsoft Purview, Unity Catalog, AWS Lake Formation, Metadata Management, Data Lineage, Data Catalog, RBAC, RLS, Column-Level Security, Encryption, Audit Controls
Data Quality: Great Expectations, DLT Expectations, Profiling, Anomaly Detection, Reconciliation, Automated Validation, Data Completeness, Accuracy and Integrity Checks
DevOps / IaC: Azure DevOps, Git, GitHub Actions, Terraform, Azure Bicep, CI/CD, GitOps, Infrastructure as Code, Policy as Code, Automated Testing, Release Automation
Visualization / BI: Power BI, DAX, Tableau, Looker, SSRS, ggplot2, Matplotlib, Semantic Models, KPI Dashboards, Incremental Refresh, Aggregations
Tools: PyCharm, Eclipse, Visual Studio, SQL*Plus, SQL Developer, TOAD, SQL Navigator, Query Analyzer, SQL Server Management Studio, Postman
Web / APIs: REST APIs, FastAPI, Flask, Apache Tomcat, Microservices
Operating Systems: Linux, UNIX, Windows
Healthcare / HIPAA: 837/835, 270/271, 276/277, EDI, Claims Data, Member Data, COB, ICD, Compliance and Secure Data Processing
Title: Senior Data Engineer
Client: State of Ohio (Department of Aging), Columbus, OH May 2023 to present.
Responsibilities:
Designed and delivered end-to-end Azure data solutions using Azure Data Factory, ADLS Gen2, Azure Synapse Analytics, Azure SQL Database, Microsoft Fabric, and Azure Databricks for enterprise analytics and AI/ML workloads.
Engineered resilient ingestion frameworks with Databricks Workflows, parameterized job clusters, PySpark, incremental loading, dependency management, retry logic, and operational controls.
Developed Snowflake ELT pipelines using dbt models, incremental models, snapshots, macros, automated tests, documentation, and reusable transformation patterns for governed analytics.
Created dimensional models across Azure Synapse, Microsoft Fabric Data Warehouse, Azure SQL Database, and Snowflake using star schema, snowflake schema, facts, dimensions, and SCD Type 2.
Built Microsoft Fabric pipelines and lakehouse solutions across OneLake with bronze, silver, and gold layers, Delta tables, schema evolution, ACID transactions, and governed transformations.
Developed advanced SQL for Azure SQL Database, Synapse, Snowflake, and analytical marts, applying CTEs, window functions, joins, stored procedures, query plans, and workload tuning.
Implemented Power BI semantic models with DAX measures, relationships, hierarchies, KPI calculations, time intelligence, incremental refresh, and self-service reporting for business users.
Established Azure Data Factory orchestration patterns for scheduled, event-driven, batch, and incremental processing while integrating Blob Storage, ADLS, SFTP, Azure SQL, and downstream warehouses.
Implemented Microsoft Purview governance with automated scans, lineage, metadata ownership, sensitivity labels, business context, and compliance tracking to improve data discoverability.
Configured RBAC, RLS, column-level security, service principals, managed identities, and least-privilege controls across Azure data services and Power BI.
Built CI/CD pipelines in Azure DevOps for ADF, Databricks notebooks, dbt assets, Fabric lakehouse components, and data models with validation, approvals, and rollback support.
Integrated AWS S3, Lambda, Glue, and Snowflake for event-driven ingestion and cross-cloud processing, supporting secure movement and transformation of enterprise datasets.
Performed source-to-target mapping, reconciliation, migration validation, UAT, cutover support, post-deployment stabilization, and production readiness assessments.
Applied advanced SQL tuning and data-engineering optimization for large analytical workloads, improving query response, resource utilization, and processing efficiency across high-volume datasets.
Implemented automated Python and SQL validation utilities to identify completeness, duplication, lineage, schema, and integration defects before production release.
Built Delta Lake pipelines with PySpark across medallion layers, supporting schema evolution, ACID transactions, replayable processing, and scalable analytical consumption.
Configured Azure Monitor and Log Analytics alerts for pipeline failures, latency, resource utilization, and operational exceptions, accelerating incident detection and recovery.
Developed semantic models and governed BI datasets that integrated Azure SQL, Synapse, Fabric, Databricks, and Snowflake sources for enterprise reporting.
Collaborated with MDM, business, analytics, and engineering teams to establish customer and product data standards, KPI definitions, ownership, and quality expectations.
Executed data profiling, validation, and reconciliation across source and target systems during cloud modernization initiatives, ensuring accuracy and business continuity.
Documented ER diagrams, data dictionaries, lineage views, pipeline dependencies, transformation rules, operational procedures, and governance artifacts for auditability.
Worked extensively in Linux/UNIX environments using Bash and shell scripting for ETL automation, file handling, scheduling, process monitoring, and log analysis.
Supported UAT execution, release coordination, production deployment, incident resolution, root-cause analysis, and stabilization of mission-critical data pipelines.
Partnered with stakeholders to translate business requirements into scalable data models, transformation specifications, governance controls, and actionable analytics solutions.
Improved enterprise analytics performance through optimized data layouts, partitioning, caching, dimensional design, and efficient transformation logic while maintaining security and compliance.
Environment: SQL Server Management Studio 2016, Visual Studio 2015, VSTS, Power BI, PowerShell, .Net,
SSIS, DataGrid, ETL Extract Transformation and Load, Business Intelligence (BI), Python, shell scripting.


Title: Senior Data Engineer
Client: Nationwide, Columbus, OH Jan 2022 to May 2023
Responsibilities:
Designed scalable ETL/ELT pipelines using Azure Data Factory, Azure Databricks, Apache Spark, and PySpark to ingest and transform petabyte-scale datasets into ADLS Gen2 and Delta Lake.
Built batch and near-real-time workflows with Delta Live Tables, Structured Streaming, Spark SQL, and event-driven processing for high-volume analytical use cases.
Developed reusable ELT transformation frameworks and enterprise data models supporting reporting, analytics, migration programs, and downstream consumption patterns.
Participated in Teradata-to-Snowflake modernization initiatives, preparing source mappings, transformation logic, validation controls, and migration readiness activities.
Implemented Azure Databricks SQL solutions with Unity Catalog, Spark SQL, Delta queries, materialized views, partitioning, Z-ordering, liquid clustering, and workload optimization.
Governed ADLS Gen2 data lakes through hierarchical namespace, partition design, lifecycle policies, encryption, access controls, and storage optimization strategies.
Automated Azure infrastructure using Terraform and Bicep, provisioning VNets, subnets, NSGs, Private Endpoints, Databricks workspaces, IAM roles, and service principals.
Established GitOps and CI/CD delivery with Azure DevOps, GitHub Actions, Terraform modules, automated testing, security scanning, policy controls, and environment promotion.
Led Azure data platform modernization integrating ADLS Gen2, Databricks, Unity Catalog, Event Hubs, Synapse Analytics, and secure private networking for governed analytics.
Created monitoring and observability with Azure Monitor, Log Analytics, Application Insights, Databricks monitoring, custom dashboards, alerts, and automated operational runbooks.
Optimized Spark and Delta Lake workloads through partition strategy, Z-ordering, liquid clustering, Photon, autoscaling, caching, and execution-plan analysis to reduce runtime and compute consumption.
Implemented data-quality controls using Great Expectations, DLT expectations, PySpark, and SQL validation rules covering profiling, anomaly detection, completeness, and accuracy.
Designed Power BI semantic models connected to Databricks and Delta Lake using star schema, relationships, DAX measures, aggregations, and governed self-service analytics.
Improved Power BI report performance through incremental refresh, aggregation design, query diagnostics, and DAX optimization for large analytical datasets.
Collaborated with business analysts and technical stakeholders to translate requirements into KPI definitions, data products, semantic layers, and production reporting solutions.
Implemented enterprise governance using Unity Catalog and Azure Purview for metadata management, lineage, profiling, access control, quality enforcement, and compliance.
Engineered reliable orchestration with ADF, Databricks Workflows, and DLT, incorporating scheduling, dependencies, retries, exception handling, and operational recovery patterns.
Built event-driven processing with Azure Functions, Data Factory triggers, Blob Storage events, Event Hubs, and Databricks Jobs to improve data availability for downstream analytics.
Used Delta Lake and Iceberg-compatible patterns for schema evolution, ACID transactions, versioning, time travel, and interoperable data access across analytics engines.
Automated identity and access operations using Python, PowerShell, and Azure SDKs for service-principal lifecycle, secret rotation, permissions, and compliance controls.
Conducted source-to-target mapping, reconciliation, migration validation, UAT support, production deployment, cutover planning, and post-migration stabilization for enterprise workloads.
Created Spark SQL and PySpark transformations for complex business logic, large-scale joins, aggregations, incremental processing, and analytical data preparation.
Partnered with governance and data stewardship teams to define metadata standards, lineage requirements, quality SLAs, ownership models, and secure access policies.
Supported incident management and root-cause analysis for pipeline failures, data-quality exceptions, performance bottlenecks, and production integration issues.
Contributed to cloud-native engineering practices across Azure storage, compute, networking, security, orchestration, analytics, and machine-learning data preparation.
Environment: Azure Databricks (Spark, PySpark, Delta Lake, Delta Live Tables, Unity Catalog), Azure Data
Factory, Azure Data Lake Storage Gen2 (ADLS Gen2), Power BI, Azure Event Hubs, Azure Service Bus,
Azure Cosmos DB, Azure Machine Learning




Title: Data Engineer
Client: Optum, Hyderabad, India Aug 2020 to Aug 2021
Responsibilities:

Designed and optimized ETL pipelines with Hive and SQL, reducing data processing latency by 35% for large-scale analytics.
Developed complex Hive queries with joins, subqueries, and window functions to process petabyte-scale data in Hadoop HDFS.
Built ingestion frameworks using Sqoop to transfer data from Oracle, SQL Server, and MySQL into Hadoop HDFS, ensuring data integrity.
Engineered Hive-based ETL processes to transform data from HDFS and ORC/Parquet files into data marts and warehouses.
Developed and maintained Informatica PowerCenter mappings, workflows, sessions, and transformations supporting enterprise ETL and data integration processes.
Collaborated with business and data teams to design source-to-target mappings and implement high-performance ETL solutions using Informatica PowerCenter.
Performed ETL testing, troubleshooting, workflow optimization, and production support activities for large-scale data warehouse environments.
Experience working with Linux-based environments for Apache Spark, Airflow, and Python applications in cloud and distributed systems.
Managed file transfers, permissions, and server configurations in Linux environments while ensuring data security and system stability.
Developed SQL scripts for data extraction, transformation, and loading, ensuring high-quality data for reporting and analytics.
Managed relational databases (Oracle, SQL Server, MySQL) for Sqoop-based ingestion into Hadoop ecosystems.
Built and maintained ETL frameworks for healthcare claims and member data using PyArrow, Pandas, and DuckDB for preprocessing and aggregation.
Migrated legacy Hive tables to Apache Iceberg format, enabling schema evolution and time-travel capabilities for compliance audits.
Implemented snapshot management and compaction strategies to maintain optimal query performance and storage efficiency.
Integrated data workflows with Flask APIs to serve analytical results to internal applications with minimal latency.
Leveraged AWS S3-based data lake storage with Iceberg catalog to support batch and incremental data ingestion.
Designed and optimized SQL queries and stored procedures in Oracle and SQL Server, improving query performance by 40%.
Created Hive-based dashboards integrated with Tableau, delivering actionable insights to non-technical stakeholders.
Optimized HDFS storage with Hive partitioning and compression, reducing storage costs by 20% in Hadoop clusters.
Developed Java utilities for data cleansing and enrichment, improving data quality in Hive and SQL-based pipelines.
Monitored Hadoop cluster health with Cloudera Manager, resolving data pipeline issues 30% faster using custom alerts.
Automated on-premises-to-Hadoop data migrations using Sqoop and Java, ensuring seamless integration with HDFS.
Delivered training on Hive, Sqoop, and Java-based ETL tools, boosting team productivity by 20% and driving best practices.

Environment: ETL, Hadoop, Hive, SQL, Java, RESTAPI.

Title: Associate Software Engineer
Colruyt Group, Hyderabad, India Mar 2019 to Aug 2020
Responsibilities:
Designed and implemented enterprise data governance framework using AWS Lake Formation and AWS Glue Data Catalog, establishing centralized metadata management, data lineage, business glossary, access policies, and data quality rules across the AWS Data Lake.
Led metadata management initiatives by configuring AWS Glue Data Catalog and AWS Lake Formation for schema management, tagging, ownership, and automated lineage tracking, enabling improved data discoverability and trust across the organization.
Developed robust Python and ScalaSpark scripts for data transformation, cleansing, enrichment, and custom business logic on Amazon EMR and AWS Glue, significantly improving pipeline maintainability and reusability.
Spearheaded the migration of Tableau Prep workflows to modern, scalable data preparation pipelines using AWS Glue and Amazon EMR, resulting in better performance, governance, and integration with the enterprise data lake.
Led end-to-end Data Warehouse modernization project, migrating from legacy on-premises systems to a governed AWS Data Lakehouse (Amazon S3 + Delta Lake / Iceberg on EMR) for improved scalability, performance, and cost efficiency.
Successfully executed Oracle databases to Data Warehouse migration initiatives, including schema conversion, data mapping, incremental ETL development, validation, and cutover using AWS Glue, AWS DMS, and Amazon EMR for multiple business-critical systems.
Designed and optimized dimensional modeling (Star/Snowflake schema) and Slowly Changing Dimensions (SCD Type 1 & 2) using Delta Lake / Apache Iceberg on Amazon S3 to support enterprise data warehousing requirements with ACID compliance.
Automated data quality validation, profiling, and monitoring processes using PySpark, Great Expectations, and AWS Glue Data Quality as part of the governance framework.
Performed advanced data pipeline development and troubleshooting on Linux/Unix environments, including shell scripting, cron jobs, log analysis, and EMR cluster management via CLI, SSM, and SSH.
Collaborated with data stewards, analysts, and governance teams to define and enforce data policies, standards, and quality SLAs across the AWS data platform.
Built self-service analytics enablement solutions by combining governed S3 data lakes, Amazon Athena, and Amazon QuickSight / Power BI, ensuring secure and compliant data access post-migration.
Implemented automated metadata scanning, classification, and sensitivity labeling using AWS Glue, AWS Lake Formation, and Amazon DataZone to support regulatory compliance

Environment: AWS, Python, Scala, Spark
Keywords: cprogramm cplusplus continuous integration continuous deployment artificial intelligence machine learning business intelligence sthree database information technology procedural language Ohio

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7713
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: