◆ Professional Summary
15+ years driving enterprise data transformation for global tech leaders and startups across cloud/SaaS, fintech, gaming, and consumer platforms. Expert in cloud-native architectures, real-time analytics, privacy-first pipelines, and AI-augmented data operations. Proven builder of high-trust, high-performance data teams at petabyte scale. Published writer and thought leader on data architecture, AI agents, and the intersection of ancient wisdom with modern technology.
◆ Core Competencies
Modern Data Architecture
Privacy-First Design
Consent Architecture
AI Agents & LLM Pipelines
Prompt Engineering
Data Governance (C1–C5)
Petabyte-Scale Systems
Global Team Leadership
Data Storytelling
★ Featured Experience
- Landed 1,000+ production changes across a multi-year privacy migration — sustained solo throughput normally requiring a team of 3–4 engineers
- Cleared the final blocker to a company-wide data warehouse consolidation program. My workstream covered the last ~20% of warehouse data — the portion nobody else could touch because it sat directly in the advertising revenue path. Completing it took the program from 80% to 100% coverage and enabled its formal close-out
- Executed that migration with zero revenue impact on the highest-sensitivity assets in the warehouse, under a constraint that any error would surface as ads revenue loss rather than a test failure
- Proved the pre-deployment validation stack was structurally blind — permission checks returned PASS for tables that didn't exist and never modeled runtime enforcement. Built the hard-gate guard system that closed the gap; every gate traces to a specific postmortem
- Reduced an entire failure class to an offline-checkable invariant. Verified against the full asset inventory: 81% of migrations required a namespace change, not the name change the tooling assumed
- Owned response for 5 production incidents including 2 SEV3s (one at $5.27M/day) — fix-forward, zero reverts, permission grants coordinated across 6+ owning teams
- Co-built an AI coding skill encoding the full remediation methodology: engineer ramp-up 2 weeks → 2 days, per-asset cycle time down 3–5×
- Work recognized in a company-wide announcement to the Data organization, with the difficulty of this workstream called out explicitly by senior engineering leadership
LLM Token & Cost Engineering
- Designed prompt templates with stable cacheable prefixes + minimal variable context — cut input tokens 60% across 800+ AI-assisted diffs
- Routed simple transforms to local template engines (zero API cost); reserved Claude for ambiguous multi-consumer rewrites only — eliminated 30% of unnecessary invocations
- Enforced retrieval-over-stuffing: pass 40-line function blocks, not 2,000-line files. Structured JSON output schemas cap response length to what's needed
- Capped agent retry loops at 3 iterations with pre-flight input validation — killed retry spirals burning 40x normal token budget on stale inputs
- Tracked cost per landed diff (not per API call) — optimized for first-pass success rate, driving effective cost to ~$0.80/remediation at 800+ diff scale
Strategic architect for VMware's Partner Data Platform — unifying fragmented partner data into a single governed analytics layer
Data Architecture & Platform Unification
- Designed and built the Partner Data Platform from scratch, unifying 10+ disparate sources — eliminating data silos blocking cross-functional decisions for years
- Built data catalog and governance model (Python/Confluence) cutting tribal knowledge reliance by 60% and ad-hoc requests by ~40%
- Implemented VMware's first structured data governance for partner data — classification, lineage, policy enforcement — achieving audit-ready status
- Improved data consistency by 60% through standardized naming, automated quality checks, and cross-team contracts
AI & Performance Engineering
- Pioneered an AI copilot for retrieval optimization and governance — one of VMware's earliest LLM-assisted data engineering deployments, reducing query dev time by ~30%
- Achieved 40%+ Spark processing time improvements through broadcast joins, caching, and skew management
- Established automated monitoring replacing reactive firefighting, reducing incident response from hours to minutes
Leadership & Cross-Functional Alignment
- Drove alignment across 5+ data teams via shared roadmaps, reducing duplicate efforts by ~25%
- Led workshops upskilling 30+ partner data consumers, building organizational data literacy
- Engaged director/VP-level stakeholders, securing budget — team grew from 3 to 8 engineers
Led high-impact projects across Meta's Ads, Commerce, and Privacy teams
Commerce Data Architecture
- Architected centralized DW unifying product/seller data — reducing decision-making cycle from weeks to days
- Launched Category Management Data Warehouse enabling cross-vertical analysis and measurable GMV growth
- SMB funnel redesign: 3x query performance improvement, scaling to tens of thousands of advertisers
Privacy Engineering & Compliance
- Led privacy remediation across 1,200+ SMB 2.0, Customer Journey, and BPO tables — establishing patterns adopted in Safe Ads
- Enhanced anonymization with Hive engineers, reducing audit preparation effort by ~60%
- Drove Salesforce ID deprecation across 1,000+ pipelines, eliminating vendor dependency affecting ~15% of joins
Data Platform Migration
- Migrated 1,000+ Hive pipelines to Spark, automating ~70% — cut projected 12-month migration to under 6 months
- Built chargeback/leakage dashboards (FGF, Dataswarm, Unidash) identifying previously undetected revenue leakage
- Designed Marketplace App reliability dashboard — reducing MTTR for payment incidents via real-time visibility
★ Additional Experience
- Migrated 1,000+ Hive pipelines to Spark; built chargeback representment & Marketplace reliability dashboards
- Led GDPR compliance: encrypted sensitive data into Dali storage using Hive, Python, and Pig; retired legacy sources
- 80+ Dataswarm pipelines for petabyte-scale ads: Cross Device Insights, Global Account Pipeline, Outcomes Datamart + Norms DB, Facebook Media & Live Monetization
- PII masking, Vertica→Redshift migration, 1,000+ table optimization
- Life Sciences BI, cloud migration, ETL modularization
- Incentive comp analytics, BusinessObjects XI
- Enterprise DW + 8 datamarts; ETL reduced from 18 hrs to 3 hrs; 40–50% sales lift via clickstream analytics
⚙ Technical Skills
Big Data & Cloud
SparkHiveHadoopDatabricksDelta LakeSnowflakeRedshiftAWSAzureGCP
Databases & Tools
MongoDBDynamoDBDataswarmInformaticaUnity CatalogPolymerFGFUnidashBusinessObjects
Programming
PythonSQLScala
AI & Privacy
Differential PrivacyClaude CodeAI AgentsPrompt Engineering
☞ Education
Data Engineering
Certificate
UC Santa Cruz, 2013–14
B.S. Electrical Eng.
Degree
Sardar Patel College of Eng.
Electronics & Comms
Diploma
Technical Board, Tamil Nadu
◆ Certifications & Publications
Certifications
VMware SaaS EssentialsHadoop FundamentalsNoSQL DatabasesNoSQL for SQL Professionals
Publications
Cloud ComputingAll About Big DataFuture Trends in BI
◆ Open-Source & Community
Interview Studio — a free, no-sign-up practice platform I built for the Data & AI Engineering community, independent of any employer. Runs entirely in the browser (sql.js + Pyodide / WebAssembly) — no backend, no telemetry, no paywall.
22
Data-Model Deep-Dives
End-to-end industry scenarios — ER diagrams, runnable SQL & worked examples
1,500+
Practice Questions
Real SQL / Python / Snowflake from 100+ companies, runnable in-browser
0
Backend · Sign-up · Cost
Fully client-side (WASM) — free and open to everyone
॥ AI & Thought Leadership
96+
Published Articles
Data architecture, AI agents, philosophy & ancient wisdom
18+
Sacred Text Guides
Interactive Sanskrit texts with transliteration & meanings
AI
Innovation
Claude Code skills, vibe coding methodology, AI agent architecture