Senior Software Engineer · Backend · Data · GraphQL

Backend and data platforms built for seasonal scale and operational clarity.

I architect backend and data platforms across AWS and GCP, lead federated GraphQL systems, and am extending that production discipline into RAG, agents, and AI evaluation pipelines.

Focus
GraphQL Federation · AWS/GCP · Data engineering · Applied AI
Location
Remote-friendly · Multi-timezone
  • 7+ Years engineering production systems Backend · data · distributed computing
  • Millions Of acres supported annually Seasonal agricultural modeling at continental scale
  • 78 Scholarly citations 3 publications · IEEE Senior Member

01 / Work

Selected work

Three focused stories spanning production agriculture, distributed scientific computing, and reproducible geospatial research.

Production platform · Climate LLC / Bayer Crop Science 2021 – present

Agricultural Recommendation & Modeling Platform

View ↗
REST/OpenAPIWunderGraphGraphQL FederationAWSGCPBigQueryEvent-driven systemsObservability

Problem

Recommendation and modeling teams need reliable access to agricultural models and geospatial data across continental operations, with demand that rises sharply during seasonal production windows.

Approach

Architect backend and data services across AWS and GCP, shape versioned REST contracts, lead a WunderGraph-based federated GraphQL platform, and connect BigQuery data workflows with production telemetry and operational controls.

Impact

The platform supports model execution and geospatial data delivery across millions of acres annually. Backward-compatible contracts, schema governance, and reliability practices keep high-concurrency seasonal workloads predictable and observable.

NIH-funded scientific platforms · UConn Health 2019 – 2021

VCell & BioSimulations Ecosystem

View ↗
Backend APIsSLURMHPCNATSDockerSingularityPythonJavaReproducibility

Problem

Researchers needed a dependable way to submit, coordinate, and reproduce heterogeneous biological simulations without manually managing solver environments or HPC job lifecycles.

Approach

Built SLURM-based dispatch and backend integration for simulation setup, metadata, workload control, lifecycle status, and result retrieval. Containerized heterogeneous solver environments and used messaging to coordinate long-running distributed work.

Impact

The infrastructure enabled reproducible execution across VCell, BioSimulations, and BioSimulators and contributed to two peer-reviewed Nucleic Acids Research publications.

Independent geospatial research · Open source Current

TerraFlow-Agro

View ↗
PythonGeospatial workflowsProvenanceSpatial validationSensitivity analysis

Problem

Agricultural suitability analysis is difficult to trust when preprocessing, assumptions, validation choices, and data lineage are not captured consistently.

Approach

Created a configuration-driven workflow with deterministic fingerprints, provenance manifests, spatial cross-validation, and sensitivity analysis across raster and climate inputs.

Impact

The project makes agricultural modeling experiments traceable, repeatable, and easier to evaluate across changing datasets and regions.

02 / Current focus

Applied AI Engineering

A hands-on transition from senior backend engineering into production AI systems. The work starts with raw model APIs and builds toward public RAG, agent, evaluation, serving, and observability projects.

In progress

Applied AI Engineering

Independent learning & build track · In progress

Building toward

  • RAG, embeddings, vector search, reranking, and retrieval evaluation
  • Agent loops, tool use, MCP, memory, and multi-step orchestration
  • Golden datasets, LLM-as-judge, regression suites, and prompt CI
  • Serving, streaming, latency, cost controls, safety, and AI observability

Engineering method

  • Build applied systems before abstracting them behind frameworks
  • Treat evaluations as automated tests, not a final demo step
  • Carry production habits—contracts, failure handling, and telemetry—into AI

Public proof grows with the work. Completed labs and deployed projects will graduate into full case studies as they ship.

03 / Approach

How I build

Four anchors for systems that must evolve safely, move data reliably, and remain understandable in production.

01

Federated API design

Contracts and evolution before implementation.

  • Shape REST and GraphQL contracts around clear ownership and failure modes.
  • Govern federated schemas across subgraphs and consumer boundaries.
  • Use compatibility checks and staged releases to reduce rollout risk.
02

Cloud data platforms

Operational systems and analytical data should reinforce each other.

  • Build across AWS services and GCP/BigQuery data workflows.
  • Design query and reporting patterns with performance and cost in mind.
  • Keep lineage, access patterns, and operational telemetry visible.
03

Asynchronous reliability

Variable demand needs bounded, replay-safe behavior.

  • Use batching, backpressure, checkpoints, fault isolation, and bounded retries.
  • Design workflows to survive partial failures and safe reprocessing.
  • Validate seasonal behavior with load testing and production signals.
04

Production readiness

Systems are finished when operators can understand them.

  • Connect logs, traces, dashboards, alerts, and SLOs into one diagnostic path.
  • Make runbooks, rollback paths, and failure handling part of delivery.
  • Apply the same discipline to AI evaluation, serving, safety, and cost.

Technical stack

Languages Primary
  • Python · Java · Scala
  • TypeScript · JavaScript
Backend & APIs Primary
  • REST / OpenAPI
  • WunderGraph · GraphQL Federation
  • Event-driven services · Microservices
Cloud Primary
  • AWS ECS/EKS · DynamoDB
  • SQS/SNS · IAM/SSM/ECR
  • GCP · BigQuery
Data & Messaging Primary
  • Data pipelines · Schema evolution
  • BigQuery · NATS · Pub/Sub
  • Asynchronous processing
Platform Active
  • Docker · Kubernetes · Singularity
  • SLURM · GitHub Actions · GitLab CI
Reliability Primary
  • Datadog · Splunk · OpenSearch
  • OpenTelemetry · SLOs · Runbooks

04 / Research

Research & recognition

Peer-reviewed work and professional service that reinforce the engineering practice—not compete with it.

78 citations h-index 3 · i10-index 3
July 3, 2026

Peer-reviewed publications

B. Shaikh, L. P. Smith, D. Vasilescu, G. Marupilla, et al.

BioSimulators: a central registry of simulation engines and services for recommending specific tools.

Nucleic Acids Research, 2022. DOI ↗

B. Shaikh, G. Marupilla, M. Wilson, et al.

RunBioSimulations: an extensible web application that simulates a wide range of computational modeling frameworks, algorithms, and formats.

Nucleic Acids Research, 2021. DOI ↗

M. A. Teja (Marupilla Akhil Teja), et al.

Analysis of exhaust manifold using computational fluid dynamics.

Fluid Mechanics: Open Access, 2016. DOI ↗

Current manuscripts

Stress-Conditioned Lightweight Flash-Drought Detection in the US Corn Belt Using MODIS NDVI and ERA5-Land

Submitted · preprint available

Cross-Validation Protocol Drives a 20-Point R² Gap in In-Season County Corn Yield Forecasting

In preparation

05 / Contact

Ready to build
the next reliable platform.

Open to senior backend and data-platform roles, GraphQL platform work, and opportunities that connect production engineering with applied AI. If the system needs to remain dependable under real operational pressure, let's talk.

Send an email