Varun Joshi

Senior Data Engineer at Amazon Web services| Helping startups build AI on data that holds up.

Seattle, United States (-08:00 UTC) English, Hindifrom Seattle, United States
Usually responds in 5 hours
Free
Price per hour
15 min30 min
Time Blocks Available
5.00
2 reviews / 2 sessions
Thu
17
Next availability

Bio

I help early-stage startups and growing organizations turn messy, scattered data into a foundation they can actually build AI on. My focus is end-to-end: architecture and pipelines, data management and quality, governance, and security — the full stack that determines whether an AI initiative succeeds or quietly stalls on bad inputs. With over 10 years in data engineering and leadership roles at Cigna, Premera and Amazon, I've seen the same pattern repeat: teams rush to adopt AI before their data is trustworthy, accessible, or compliant. That gap is where I work. I advise founders and data leaders on building pipelines that scale, governance frameworks that don't slow teams down, and security practices that hold up under real scrutiny — not just checkbox compliance. What I help with: Data architecture and pipeline design built for scale from day one, not retrofitted after the first outage. Data governance and quality frameworks that make "AI-ready data" a practical standard rather than a slogan. Security and compliance advisory tailored to early-stage constraints — resource-conscious, risk-aware, audit-ready. Hands-on guidance for teams standing up their first data platform or maturing an existing one ahead of AI adoption. My approach is practical over theoretical. Startups don't have the headcount or runway for enterprise-grade complexity, so I focus on what's defensible now and what scales later — right-sized governance, pragmatic security postures, and architectures that won't need to be torn down at Series B. I work best with founders and technical leaders who know data quality is the bottleneck for their AI ambitions and want a second set of eyes that's built and governed data systems before, not just theorized about them.

Expertise


  • Artificial intelligence

    Varun has Masters in Informations systems with Data Science as his major along with experience of over 12 years in Data engineering.Varun is passionate about leveraging Artificial Intelligence (AI) and Machine Learning (ML) and applying thoughtful mechanisms to multiple projects.

  • Career guidance

    Career guidance is where I work directly with individuals rather than organizations — engineers, analysts, and early-career data professionals trying to figure out where to grow next in a field that's shifting fast under AI. I help people map out realistic paths: whether to go deep on data engineering versus broaden into analytics or applied AI, how to build a portfolio that actually demonstrates judgment rather than just tool familiarity, and how to read the market so you're optimizing skills

  • Data science

    Varun has Masters in Informations systems with Data Science as his major along with experience of over 12 years in Data engineering.Varun is passionate about leveraging Artificial Intelligence (AI) and Machine Learning (ML) and applying thoughtful mechanisms to multiple projects.

  • Leadership

    I help technical leads step into leadership: owning roadmaps, making calls under uncertainty, and turning deep data expertise into decisions their teams actually trust and follow.

  • Product analytics

    Product analytics is where I help startups turn raw usage data into a trustworthy signal for both business decisions and AI. As products increasingly feed models and personalization features, the quality of the underlying event data determines what's actually possible downstream — funnels, retention curves, and AI features are only as good as what's tracked. I work with teams to build product analytics from the ground up.

  • Product launches

    I help founders ship AI products on solid technical and product footing, not just working demos. Before launch, I dig into MVP scope, where the model is likely to fail, how you'll evaluate it, how user feedback gets back into the loop, what metrics actually matter at launch, and where to keep a human in the loop versus automate. This is the right fit if you're heading toward an AI MVP and need a clear-eyed view of what can break, what to measure.

  • Technology and tools

    I help teams choose tools that fit their stage, not their ambitions: warehouses, pipelines, governance, and analytics stacks that scale without adding complexity you can't yet support.

Toolkit


  • AWS (Amazon Web Services) logo

    AWS (Amazon Web Services)

    7 years of experience

    Varun is a Senior Data Engineer with AWS since 2022 and has delivered numerous projects along with having experience with multiple AWS services such as Spark, S3, Redshift, EMR, Lake formation, Dynamo DB etc.

  • MySQL logo

    MySQL

    12 years of experience

    Building end to end data pipelines where SQL is used to transform the data into usable and clear metrics which can be reported by Analysts and utilized for Data science models.

  • Python logo

    Python

    9 years of experience

    Expert in utilizing python for data ingestion, data transformation and building machine learning models. Building API's and scraping data from various sources.

Experience

  • AWS

    Senior Data Engineer
    aws.amazon.com/

    Led end-to-end design of a large-scale financial data migration, consolidating multiple regional invoice sources into a unified immutable double-entry accounting system; collaborated across 4+ teams as primary technical SME, resolving critical data discrepancies ahead of cutoff deadlines with zero production issues. Architected an intelligent Data Quality & Anomaly Detection framework by benchmarking ML algorithms (Isolation Forest, Local Outlier Factor, k-NN) and DQ libraries (AWS Glue DQ, Great Expectations, Pydeequ), establishing automated validation pipelines across cloud storage and warehouse layers. Redesigned a critical revenue metric (Total Contract Value) to incorporate agreement lifecycle events; engineered an automated comparison tool that replaced 4–5 days of manual analysis with a sub-10-minute script, enabling sales planning workflows for cross-functional stakeholders. Architected a dedicated reporting cluster with encryption, secrets management, data sharing, and WLM queue configuration; led a zero-downtime parallel migration of 40+ reporting pipelines, eliminating a single point of failure that had caused multi-day business outages. Spearheaded a new product taxonomy rollout across 50+ ETL scripts spanning renewals, marketing, insights, and revenue domains; conducted phased deployments with thorough impact analysis and stakeholder communications, ensuring zero disruption to downstream consumers.

  • Premera Blue Cross

    Principal Data Engineer
    premera.com/visitor?region=pbcwa

    Served as lead engineer for orchestrating and maintaining critical monthly data exchange between Premera and Blue Cross Blue Shield, ensuring reliability and compliance across organizational boundaries. This initiative, compliant with federal HIPAA/CMS interoperability guidelines, also supports electronic data interchange (EDI) for faster provider claim processing and eventually higher customer satisfaction. Drove Premera's data platform modernization by leading a full-lifecycle Netezza-to-Snowflake migration — encompassing design, development, validation, and cutover — enabling the organization to retire legacy on-premise infrastructure and adopt cloud-native services, advanced analytics, and next-gen data engineering tools at scale. Built batch and near-real-time data pipelines using Qlik Replicate and DataStage within the Corporate Data & Analytics team; authored monitoring reports in Tableau to track production job health for QlikReplicate workflows. Conducted POC to acquire data via REST APIs using Python's Requests module and Pandas for data manipulation, demonstrating feasibility of API-driven ingestion patterns. Automated Data Validation: Developed a system that uses an LLM to automatically generate data quality rules from schema and data samples, eliminating manual rule authoring. AI Enablement & Training: Planned and facilitated an org-wide workshop training 20+ data engineers on agentic AI systems; reduced agent development and approval time from weeks to a single day.

  • Cigna

    Software Engineering Advisor

    Designed and implemented end-to-end big data pipelines on the Hortonworks stack using Apache Spark, Kafka, Spark Streaming, and HBase; developed a fault-tolerant ML data pipeline whose model delivered $1M+ in savings in its first year. Built low-latency ETL pipelines loading data into HDFS via Sqoop, Hive, and Java; leveraged Datasets and Spark SQL to re-engineer legacy stored procedures. Used advanced Sqoop features for efficient ingestion from Oracle, SQL Server, and Teradata; applied Spark performance tuning, caching, and memory management best practices. Designed and maintained a Change Data Capture pipeline using Qlik Replicate, Kafka, and Spark Streaming to load pharmacy data into HDFS in near real time. Practiced CI/CD using Git, GitHub, and Jenkins to maintain code quality and streamline deployments

Made within Glyfada