Systems & Data Infrastructure Engineer
Posted Updated
Functional Focus: Data Engineering, Real-Time Streaming, and HPC Job Orchestration
- Scientific Data Architecture: Design crash-safe, binary data formats capable of handling high-speed incremental writes from compute jobs and concurrent reads for query-time calculations.
- Operational Data Modeling: Design and maintain a relational database schema via an async ORM, backend-agnostic across database engines.
- HPC Workload Management: Manage the lifecycle of batch and interactivecompute jobs, handling subprocess monitoring, status polling, and cluster filesystem coherency.
- Real-Time API & Streaming: Build async backend services and long-lived, server-pushed event channels with backpressure handling, and enforce role-based access control on all APIs.
- AI Service Integration: Integrate and orchestrate calls to an AI/analytics service from the backend, coordinating with application state.
Requirements
Mandatory:
- Python 3, async/await
- FastAPI (or similar async Python web framework)
- Relational DB modeling, async ORM (e.g. Tortoise, SQLAlchemy), SQL
- Server-Sent Events or WebSockets, backpressure handling
- Role-based access control (RBAC), token-based auth (JWT/OAuth)
- Subprocess management, batch job lifecycle/status polling
- Binary/streaming file format design, crash-safe writes
- REST/SDK integration with an external AI/LLM service
Optional:
- Go or another async-capable backend language
- GraphQL
- PostgreSQL administration, Alembic/migrations tooling
- Message queues/brokers (Redis, RabbitMQ, Kafka)
- Air-gapped/offline deployment experience
- HPC schedulers (SLURM, PBS), Linux cluster filesystems
- NumPy/columnar formats (Parquet, HDF5)
- Tool-calling / function-calling orchestration patterns
Benefits
We offer great career growth, ESOPs, Gratuity, PF and Health Insurance.
Skills
As published by workable · 18 questions
Basics
First name, Last name, Email, Headline, Phone, Address, Photo, Education, Experience, Summary, Resume, Cover letter
Short answers (3)
- What is your current CTC?
- What is your expected CTC?
- In how many days you can join?
Pick from a list (15)
- Have you built a system that writes data incrementally at high speed while other processes concurrently read from it?
- Have you designed a relational database schema for an operational (not just analytical) system?
- Have you used an async ORM (e.g. SQLAlchemy async, Tortoise ORM) and written database code that is backend-agnostic (works across more than one database engine)?
- Have you written asynchronous Python (async/await) in a production backend service?
- Have you built REST APIs with a Python async framework (e.g. FastAPI or similar)?
- Have you implemented a long-lived, server-pushed communication channel (Server-Sent Events or WebSockets) and handled backpressure when the client can't keep up?
- Have you implemented role-based access control (RBAC) or token-based authentication (JWT/OAuth) on an API?
- Have you managed the lifecycle of long-running background jobs — launching subprocesses, polling status, handling failures/timeouts?
- Have you worked with a shared/cluster filesystem where multiple processes read and write concurrently, and had to reason about consistency?
- Have you integrated a backend service with an external AI/LLM service via REST or an SDK, and coordinated the response with application state?
- Have you designed and run database schema migrations in a live system (e.g. with Alembic or similar)?
- Have you worked with a message queue or broker (e.g. Redis, RabbitMQ, Kafka) in a production system?
- Have you deployed or operated a system in an air-gapped or fully offline environment?
- Have you used an HPC job scheduler (e.g. SLURM, PBS) or worked with columnar/scientific data formats (e.g. Parquet, HDF5, NumPy arrays) at scale?
- Are you willing to relocate to Bangalore?