Career

Nine years across data platforms, databases, and production agentic systems.

Nine years across data platforms, database lifecycle operations, system architecture, and agentic systems in production. Roles, measured outcomes, and the stack behind each.

Download the PDF Read the case studies Email Basant

15 TB+ production data footprint
100M+ events processed daily
99.9% platform uptime sustained
40% lower infrastructure cost

Experience at a glance

Role Period Scope
Senior Data Engineer, UXCam 2024 – present Owns platform reliability and governance, the database layer beneath it, and the production agent workflows; architecture review and mentoring
Data Engineer, UXCam 2020 – 2024 Built the batch and streaming processing, the lakehouse models, and the multi-engine serving layer
Project Leader, SVCET 2019 – 2020 Planning and delivery of an NLP dialogue-system project with a small team
Software Developer, SV Technology 2017 – 2019 Python backend services, PostgreSQL schema work, reporting, and CI

Senior Data Engineer · UXCam

February 2024 – present · Working with a global team

  • Own reliability and governance for a 15 TB+ platform processing 100M+ events daily, with 99.9% uptime, explicit service objectives, incident-ready runbooks, and reviewable data contracts.
  • Operate the database layer the platform depends on: clustering and replication, backup and recovery, access control, query and index tuning, and upgrade windows across SQL and NoSQL engines.
  • Design and operate the production agent workflows — LangGraph and Google ADK orchestration, MCP tool servers, retrieval and evaluation — wired into the existing platform so model output lands as typed, provenanced rows rather than a transcript. Analytics delivery 60% faster with 75% less analyst effort, with review and fallback paths kept in place.
  • Lead architecture and schema reviews before new stores, pipelines, and agent workflows reach production; mentor 5 engineers through pairing, code review, and written design feedback; team delivery time improved 30%.

Data Engineer · UXCam

February 2020 – February 2024

  • Rebuilt core Spark/PySpark processing for 5 TB+ per day, improving processing time by 50%.
  • Administered distributed SQL and NoSQL stores used for serving and product state—modeling, partitioning, replication, and query plans rather than treating the database as a black box.
  • Made schema evolution, replay, late data, and backfills designed interfaces with idempotent loads and explicit watermarks.
  • Reduced p95 query latency by 50% across 100M+ analytical queries per day through partitioning, clustering, and model redesign.
  • Led a Kubernetes migration of the data and database infrastructure that lowered cost by 40% while sustaining 99.9% uptime.

Capability map

  1. Data platforms

    Batch and streaming ingestion, Spark/PySpark processing, Airflow orchestration, Iceberg and dbt modeling, Kubernetes capacity, contracts, lineage, and cost. Evidence: production platform case study.

  2. Database lifecycle

    PostgreSQL, ClickHouse, Trino, Citus, TimescaleDB, Redis, MongoDB, and MySQL/MariaDB across modeling, HA, replication, backup and PITR, upgrades, access control, and query performance. Evidence: measured failover lab.

  3. Agentic systems

    Production agents on LangGraph, LangChain, Google ADK and MCP tool servers: bounded tools, retrieval, evaluation before rollout, typed output, and a fallback path that fails honestly. Built against the platform already running. Evidence: governed AI case study.

  4. System architecture

    Application boundaries, schema design, dual OLTP/OLAP stores, typed contracts, and recovery paths a product can actually replay. FastAPI, PostgreSQL, ClickHouse, and the service layers between them. Evidence: independent product case study.

The stack

The stack

Tiers are self-assessed and deliberately conservative. Tier 1 means sustained production use; Tier 2 means meaningful delivery on a narrower surface; Tier 3 marks working knowledge from evaluation, prototyping, or lab work.

Technologies I have operated, debugged, and reviewed as part of long-lived systems.

  • Python since 2017

    Backend services, data platforms, automation, and AI tooling

  • SQL since 2017

    Transactional, analytical, and reliability work — PostgreSQL tutorial

  • Apache Spark / PySpark since 2020

    Batch and streaming data processing — medallion pipeline

  • Apache Kafka since 2020

    Event-driven and streaming systems

  • Apache Airflow since 2020

    Workflow orchestration, replay, and recovery design

  • Apache Iceberg since 2022

    Table evolution, maintenance, and data modeling — compaction and rollback, measured

  • dbt since 2021

    Transformation contracts, testing, and documentation

  • PostgreSQL since 2017

    Application schemas, replication, HA, backup, and upgrades — failover lab

  • ClickHouse since 2021

    Analytical modeling and operational practice

  • Trino since 2021

    Federated SQL and query modeling — Trino on Iceberg via Polaris

  • Kubernetes since 2022

    Container orchestration and data workloads

  • LangGraph / LangChain since 2024

    Production agent orchestration and tool use — agent case study

  • Google ADK since 2024

    Agent applications on an OSS framework, integrated into product workflows — agent case study

  • MCP since 2024

    Tool servers and apps integrated into existing systems — agent case study

  • Prometheus + Grafana since 2021

    Service and data-reliability signals — Reliability lab

Used to deliver real systems, with less breadth or duration than Tier 1.

Agentic systems
CrewAI · FastMCP · Milvus · MLflow · Pydantic
Data platform
Databricks + Unity Catalog · Amazon Kinesis · Spark Structured Streaming · Debezium / CDC · Delta Lake
Databases
Redis · MongoDB · Citus · TimescaleDB · MySQL / MariaDB · Patroni / pgBackRest
Product & infra
FastAPI · Celery · Next.js · Docker · Ansible · Terraform · GitHub Actions · Helm

Evaluated, prototyped, or run in a lab; listed to show where the depth stops.

Apache Flink · Apache Beam · Prefect · LlamaIndex · AutoGen · Haystack · Semantic Kernel · DSPy · OpenAI Agents SDK · Pinecone · pgvector · Chroma · FAISS · Weaviate · ScyllaDB · MariaDB + Galera · SolrCloud · Elasticsearch · Apache Druid · Snowflake · BigQuery · Amazon Redshift · DuckDB · Hadoop · Apache Hive · Scala · Java · Azure Synapse · Azure Data Factory · Power BI · Tableau · Looker · Superset · Metabase · PyTorch · Polars · PyArrow · Jenkins · LangSmith · Langfuse · Opik · OpenTelemetry · Datadog

Education and languages

Bachelor of Technology, Computer Science and Engineering — JNTUA College of Engineering, Anantapur, India, 2015–2019.

English: C1 · Nepali: native

Contact

technobasant9@gmail.com · LinkedIn · GitHub