Software Development Engineer II  ·  Equity Data Science  ·  Mumbai, IN

Nikhil Sai

I build web services, then spend my time making sure they don't fall over.

Résumé Email GitHub LinkedIn
Fig. 01 — load shedder, live traffic burst

request rate shedding readiness

Schematic, drawn from the numbers in the incident write-up — not a live feed. When a pool's acquire queue hit 2× pool size, doomed requests got an immediate 503 + Retry-After instead of a place in a queue they'd never reach the front of. Across the burst: ~2,500 requests shed over 97 paired engage/disengage cycles of 59–365 ms each. Zero readiness failures, no restarts, nobody paged.
AboutMumbai · hybrid

I work on the part that starts after the feature works

Backend engineer, coming up on three years of production ownership at an investment-research platform. Most of what I ship lives in the unglamorous middle of a service: connection pools, readiness semantics, retry and idempotency rules, and the code that decides what happens when traffic arrives faster than the database can answer.

The failures I find most interesting are the quiet ones — a connection that is never returned, a health check that says PASS on a task that is wedged, a queue that accepts work it can never finish. They don't page anyone until they've taken the whole box down. So I spend my time bounding the blast radius and making recovery something the system performs on its own at 3 a.m., rather than something I perform.

Nikhil Sai, seated cross-legged with a laptop.
Nikhil Sai Allam
B.Tech ECE, PDPM IIITDM Jabalpur
Resilience under load
Pool accounting, load shedding with hysteresis, health-check semantics, graceful degradation.
Event-driven backends
Kafka partitioning and ordering, idempotency keys, dead-letter replay, distributed locks.
Operability
OpenTelemetry and Datadog, incident write-ups, architecture decision records, integration tests over the async paths.
Work4 roles · 2021 → now

Where I've done it

Jan 2024 — now2 yrs 9 mos

Equity Data Science — Software Development Engineer II

Mumbai · hybrid · investment research & portfolio analytics

  • Own the reliability of a Node.js / TypeORM service on Aurora PostgreSQL behind RDS Proxy — the three incidents in field notes are all from this system.
  • Built the secure portfolio-ingestion path: presigned S3 uploads, server-side validation and transformation, and error diagnostics a client can act on without calling support.
  • Shipped automated report delivery on Step Functions and Fargate, driving headless Puppeteer across the analytics dashboards and mailing the output to clients.
  • Built internal Angular / Node / GraphQL tooling for AWS Batch — job configuration, execution tracking, status, and controlled re-runs.

Jan — Jun 20236 mos

Amazon — Software Development Engineer, Intern

Hyderabad · hybrid

  • Stood up stress-testing infrastructure in CloudFormation and TypeScript, generating production-scale traffic of up to 1,000 TPS per host to surface bottlenecks before customers did.
  • Onboarded clients across the FE and NA regions with region-specific infrastructure, and led the migration off a deprecated address-upload client.
  • Cut alarm noise by routing failed messages to S3 instead of a DLQ, reducing Sev-2 incidents.

Jan — Apr 20224 mos

Logicwind — Full Stack Developer

Jabalpur · contract

  • Shipped user-facing features in Next.js for production web applications.

Oct — Dec 20213 mos

HIDS Technologies — Full Stack Developer

Remote · contract

  • Delivered a MERN application in a three-month window, with real-time availability tracking and CI/CD pipelines behind it.
Field notes3 incidents · one service

Three things that broke

All three come from the same service: Node.js and TypeORM on Aurora PostgreSQL, behind an RDS Proxy, fronted by ECS. They ran in sequence — each fix exposed the next problem underneath it.

Connections that never came back

Resolved
Symptom
Recurring instance-wide outages. Every request on the box would sit and wait for a database connection that was never going to arrive.
Cause
Streaming NDJSON exports of 100,000+ rows could release their pooled connection twice — or not at all. A client that aborted mid-stream left the slot checked out with a dirty session. Every one of those was a permanent, silent leak, and they accumulated until the pool was gone.
Change
Idempotent release guards so the connection is returned exactly once, in a clean session state. A five-second bounded fallback releaser with retries behind it, and a hard destroy rather than a return whenever the state is uncertain. Transaction timeouts underneath the whole thing.
Result
Permanent leaks became bounded, self-healing recovery — ~6 s in the worst case.

Busy is not the same as dead

Health semantics
Symptom
ECS was recycling healthy-but-saturated tasks while keeping genuinely wedged ones in rotation. The readiness check could not tell the two states apart, so it got both calls wrong.
Change
Rebuilt the probe: concurrent master and replica checks on a three-second deadline, explicit classification of pool state, and two-strike hysteresis before a verdict is allowed to flip.
Result
Dead tasks are recycled within ~2 probe intervals. Saturated tasks stay in service and drain. No more false-healthy tasks, and no more churn for churn's sake.

Shedding on purpose

Load shedding
Idea
A request that will time out anyway is worse than a rejected one — it holds a slot, deepens the queue, and still fails. Better to say no immediately and say when to come back.
Change
A per-pool, method-gated load shedder. Once a pool's acquire queue reaches twice the pool size, new work gets a 503 with a Retry-After header rather than a place in line. Recovery is hysteresis-gated so it can't flap between states.
Result
In a live burst it shed ~2,500 requests across 97 perfectly paired engage/disengage cycles of 59–365 ms — zero readiness failures, no restarts, no human intervention.

Portfolio ingestion

Presigned S3 uploads, multithreaded parsing and validation, and errors that name the row and the reason.

~60% faster on large files

Report delivery

Step Functions and Fargate orchestrating headless Puppeteer over date-specific dashboard views, emailed on completion.

1,000+ reports/day · ~90% less manual work

Batch console

Internal Angular, Node and GraphQL tooling to configure, launch, track and re-run AWS Batch jobs safely.

Operator-facing · used daily
Building now2 active · Java 25

What I'm building now

4bid

Real-time auction platform

Jun 2026 → now · source private

  • Java 25
  • Spring Boot
  • Kafka
  • PostgreSQL
  • Redis
  • Angular
  • Testcontainers
  • A bid returns HTTP 202 with a pre-generated bid ID, then settles asynchronously through Kafka partitioned by auction ID into a row-locked PostgreSQL write. Idempotency keys make redeliveries no-ops, and per-auction ordering holds under concurrency.
  • Fault tolerance is explicit: Resilience4j circuit breakers, retry with backoff, a dead-letter queue with an operator replay endpoint, and a Redis distributed lock that settles each expired auction exactly once across instances.
  • Live updates fan out over Redis Pub/Sub to a Node.js WebSocket gateway and degrade to REST reconciliation — a real-time failure never fails a bid.
  • Exposed to LLM clients through an MCP server with 14 tools and an embedded chat widget. 11 ADRs written; the async paths are covered by Testcontainers integration tests.
WRITE PATH ClientPOST /bids API202 + bid id Kafkakey: auctionId Workeridempotent Postgresrow lock LIVE PATH — DEGRADES TO REST RECONCILIATION SettlementRedis lock Redispub/sub WS gatewayNode.js Clientlive updates

Pulse

Self-hosted observability platform

Jul 2026 → now · source private

  • Java 25
  • Spring Boot
  • Kafka
  • ClickHouse
  • Redis
  • OpenTelemetry
  • Ingest is standardised on OpenTelemetry, so any application in any language integrates by repointing its exporter. No platform-specific SDK required — the first-party SDKs stay thin configuration wrappers rather than proprietary instrumentation.
  • The ingest spine keeps storage I/O off the hot path: authenticated OTLP requests buffer into Kafka partitioned by application ID, and at-least-once consumers batch-insert into ClickHouse with TTL retention and pre-aggregated rollups. A storage outage costs ingest lag, not telemetry.
  • Control plane: JWT operator auth, application registration issuing SHA-256-hashed API keys with dual-key rotation grace windows, and per-app event quotas and retention overrides.
  • Roadmap sequenced logs → traces → metrics → live tail → alerting, with every phase gated on its own acceptance test.
INGEST SPINE — A STORAGE OUTAGE COSTS LAG, NOT TELEMETRY Your appOTLP exporter Ingest APIAPI key auth Kafkakey: appId Consumerbatch insert ClickHouseTTL + rollups
Stackwhat I reach for

The tools

LanguagesJava, TypeScript, JavaScript, Python, SQL, C/C++
BackendNode.js, Spring Boot, NestJS, Express, GraphQL, REST
MessagingKafka, SQS & SNS, Redis Pub/Sub
DataPostgreSQL (Aurora, RDS Proxy), ClickHouse, Redis, DynamoDB, MongoDB, Elasticsearch
FrontendAngular, React, Next.js, Nx
Cloud & opsAWS — ECS, Fargate, Lambda, S3, Step Functions, Batch · CloudFormation, Docker, GitHub Actions
ObservabilityOpenTelemetry, Datadog
PracticeDistributed systems design, TDD with Testcontainers, ADRs, incident write-ups, data structures & algorithms
EducationPDPM IIITDM Jabalpur — B.Tech, Electronics & Communication Engineering, 2019–2023 · GPA 8.0/10
Contactreplies within a day

Get in touch

Happy to talk about Kafka, database internals, or why your health check is lying to you. If you're hiring for backend or platform work — or you just want a second pair of eyes on something that keeps falling over — send me a note.