Download the PDF Read the case studies Email Basant
Experience at a glance
| Role | Period | Scope |
|---|---|---|
| Senior Data Engineer, UXCam | 2024 – present | Owns platform reliability and governance, the database layer beneath it, and the production agent workflows; architecture review and mentoring |
| Data Engineer, UXCam | 2020 – 2024 | Built the batch and streaming processing, the lakehouse models, and the multi-engine serving layer |
| Project Leader, SVCET | 2019 – 2020 | Planning and delivery of an NLP dialogue-system project with a small team |
| Software Developer, SV Technology | 2017 – 2019 | Python backend services, PostgreSQL schema work, reporting, and CI |
Senior Data Engineer · UXCam
February 2024 – present · Working with a global team
- Own reliability and governance for a 15 TB+ platform processing 100M+ events daily, with 99.9% uptime, explicit service objectives, incident-ready runbooks, and reviewable data contracts.
- Operate the database layer the platform depends on: clustering and replication, backup and recovery, access control, query and index tuning, and upgrade windows across SQL and NoSQL engines.
- Design and operate the production agent workflows — LangGraph and Google ADK orchestration, MCP tool servers, retrieval and evaluation — wired into the existing platform so model output lands as typed, provenanced rows rather than a transcript. Analytics delivery 60% faster with 75% less analyst effort, with review and fallback paths kept in place.
- Lead architecture and schema reviews before new stores, pipelines, and agent workflows reach production; mentor 5 engineers through pairing, code review, and written design feedback; team delivery time improved 30%.
Data Engineer · UXCam
February 2020 – February 2024
- Rebuilt core Spark/PySpark processing for 5 TB+ per day, improving processing time by 50%.
- Administered distributed SQL and NoSQL stores used for serving and product state—modeling, partitioning, replication, and query plans rather than treating the database as a black box.
- Made schema evolution, replay, late data, and backfills designed interfaces with idempotent loads and explicit watermarks.
- Reduced p95 query latency by 50% across 100M+ analytical queries per day through partitioning, clustering, and model redesign.
- Led a Kubernetes migration of the data and database infrastructure that lowered cost by 40% while sustaining 99.9% uptime.
Capability map
-
Data platforms
Batch and streaming ingestion, Spark/PySpark processing, Airflow orchestration, Iceberg and dbt modeling, Kubernetes capacity, contracts, lineage, and cost. Evidence: production platform case study.
-
Database lifecycle
PostgreSQL, ClickHouse, Trino, Citus, TimescaleDB, Redis, MongoDB, and MySQL/MariaDB across modeling, HA, replication, backup and PITR, upgrades, access control, and query performance. Evidence: measured failover lab.
-
Agentic systems
Production agents on LangGraph, LangChain, Google ADK and MCP tool servers: bounded tools, retrieval, evaluation before rollout, typed output, and a fallback path that fails honestly. Built against the platform already running. Evidence: governed AI case study.
-
System architecture
Application boundaries, schema design, dual OLTP/OLAP stores, typed contracts, and recovery paths a product can actually replay. FastAPI, PostgreSQL, ClickHouse, and the service layers between them. Evidence: independent product case study.
The stack
The stack
Tiers are self-assessed and deliberately conservative. Tier 1 means sustained production use; Tier 2 means meaningful delivery on a narrower surface; Tier 3 marks working knowledge from evaluation, prototyping, or lab work.
Sustained production experience
Technologies I have operated, debugged, and reviewed as part of long-lived systems.
-
Python since 2017
Backend services, data platforms, automation, and AI tooling
-
SQL since 2017
Transactional, analytical, and reliability work — PostgreSQL tutorial
-
Apache Spark / PySpark since 2020
Batch and streaming data processing — medallion pipeline
-
Apache Kafka since 2020
Event-driven and streaming systems
-
Apache Airflow since 2020
Workflow orchestration, replay, and recovery design
-
Apache Iceberg since 2022
Table evolution, maintenance, and data modeling — compaction and rollback, measured
-
dbt since 2021
Transformation contracts, testing, and documentation
-
PostgreSQL since 2017
Application schemas, replication, HA, backup, and upgrades — failover lab
-
ClickHouse since 2021
Analytical modeling and operational practice
-
Trino since 2021
Federated SQL and query modeling — Trino on Iceberg via Polaris
-
Kubernetes since 2022
Container orchestration and data workloads
-
LangGraph / LangChain since 2024
Production agent orchestration and tool use — agent case study
-
Google ADK since 2024
Agent applications on an OSS framework, integrated into product workflows — agent case study
-
MCP since 2024
Tool servers and apps integrated into existing systems — agent case study
-
Prometheus + Grafana since 2021
Service and data-reliability signals — Reliability lab
Shipped on a narrower surface
Used to deliver real systems, with less breadth or duration than Tier 1.
- Agentic systems
- CrewAI · FastMCP · Milvus · MLflow · Pydantic
- Data platform
- Databricks + Unity Catalog · Amazon Kinesis · Spark Structured Streaming · Debezium / CDC · Delta Lake
- Databases
- Redis · MongoDB · Citus · TimescaleDB · MySQL / MariaDB · Patroni / pgBackRest
- Product & infra
- FastAPI · Celery · Next.js · Docker · Ansible · Terraform · GitHub Actions · Helm
Working knowledge
Evaluated, prototyped, or run in a lab; listed to show where the depth stops.
Apache Flink · Apache Beam · Prefect · LlamaIndex · AutoGen · Haystack · Semantic Kernel · DSPy · OpenAI Agents SDK · Pinecone · pgvector · Chroma · FAISS · Weaviate · ScyllaDB · MariaDB + Galera · SolrCloud · Elasticsearch · Apache Druid · Snowflake · BigQuery · Amazon Redshift · DuckDB · Hadoop · Apache Hive · Scala · Java · Azure Synapse · Azure Data Factory · Power BI · Tableau · Looker · Superset · Metabase · PyTorch · Polars · PyArrow · Jenkins · LangSmith · Langfuse · Opik · OpenTelemetry · Datadog
Education and languages
Bachelor of Technology, Computer Science and Engineering — JNTUA College of Engineering, Anantapur, India, 2015–2019.
English: C1 · Nepali: native