Senior Database Administrator
Senior Database Administrator
Location: Bengaluru, India
Department: AI Solutions Engineering & Support
Experience: 5-8
Position Summary
Primary Responsibilities — Database Administration & Platform Engineering
Configuration, Setup & Optimization
- Provision and configure PostgreSQL, MongoDB, and Databricks environments from the ground up, applying tuned parameters, storage layouts, and connection strategies appropriate to very large data volumes.
- Optimize database configuration on an ongoing basis — memory, buffers, connection pooling, indexing strategy, partitioning/sharding, vacuum/compaction, query planning, and workload isolation — to sustain performance as volumes grow.
- Stand up schemas and configurations based on functional and non-functional requirements gathered from engineering, analytics, and product stakeholders, balancing normalization, access patterns, and performance.
- Manage Databricks clusters, workspaces, and compute policies, tuning Spark configurations, autoscaling behavior, and storage (Delta Lake) layouts for efficient processing of very large datasets.
Monitoring, Metrics & Observability
- Design and implement monitoring across all three platforms — throughput, latency, query performance, replication lag, lock contention, storage utilization, and resource saturation.
- Establish meaningful metrics, dashboards, and alerting thresholds that surface issues early and drive proactive intervention rather than reactive firefighting.
- Define and track service-level objectives (SLOs) and key operational indicators, and report on the health and performance of the data estate.
Elastic Scaling & Capacity Management
- Resize and scale environments — vertically and horizontally — based on real-time and forecasted demand, ensuring performance during peaks while controlling cost during troughs.
- Implement and tune autoscaling, sharding, replication, and partitioning strategies to accommodate sustained growth in data volume and concurrency.
- Continuously evaluate cost-to-performance trade-offs across cloud storage and compute tiers.
Secondary Responsibilities — Operations, Reliability & Analysis Support
Production Troubleshooting & Proactive Reliability
- Troubleshoot and resolve production data issues, including slow queries, data anomalies, replication problems, contention, and platform incidents — often under time pressure.
- Monitor for performance degradation and take proactive action to prevent outages before they impact users, based on data-driven signals and early warning indicators.
- Participate in incident response and on-call rotations, conduct root-cause analysis, and drive durable fixes and preventative measures.
Capacity Planning & Maintenance
- Plan for capacity upgrades, forecasting growth in storage, compute, and I/O, and scheduling upgrades ahead of demand.
- Perform routine and preventative maintenance — patching, version upgrades, index maintenance, vacuum/compaction, backup verification, and configuration hygiene — with minimal disruption.
- Maintain runbooks, operational documentation, and disaster-recovery procedures.
Data Analysis Support
- Support data analysis activities by assisting analysts and engineers with query optimization, data access, and structuring data for efficient analytical workloads.
- Partner with analytics teams to ensure Databricks and downstream platforms deliver reliable, performant access to large datasets.
Architecture & Design
- Architect high-performance database designs capable of sustaining very high transaction and query volumes without degradation.
- Design high-availability (HA) architectures — replication, clustering, failover, and multi-zone/multi-region topologies — to meet demanding uptime targets.
- Design and validate disaster-recovery (DR) strategies, including backup/restore, replication, failover runbooks, and defined RPO/RTO targets, with regular DR testing.
- Translate business and technical requirements into scalable, resilient, and cost-effective platform designs across PostgreSQL, MongoDB, and Databricks.
Required Qualifications
- Extensive hands-on database administration experience, with senior-level depth in PostgreSQL, MongoDB, and Databricks.
- Demonstrated experience scaling data environments for very large data volumes, including sharding, partitioning, replication, and performance tuning at scale.
- Deep expertise in database configuration, optimization, indexing, and query performance tuning.
- Proven experience designing and operating high-availability and disaster-recovery solutions for high-volume databases.
- Strong skills in monitoring and observability — instrumenting metrics, dashboards, and alerting to drive proactive operations.
- Experience with capacity planning, maintenance, patching, and upgrade management in production environments.
- Solid command of SQL and a scripting/automation language (e.g., Python, Bash) for operational tooling and automation.
- Working knowledge of cloud data platforms, storage/compute scaling, and Delta Lake / Spark concepts within Databricks.
- Strong troubleshooting and root-cause analysis skills, with the ability to stay calm and effective during production incidents.
Preferred Qualifications
- Experience with infrastructure-as-code and automation (e.g., Terraform, Ansible) and CI/CD for database changes.
- Familiarity with containerization and orchestration (Docker, Kubernetes) for database and data workloads.
- Exposure to data governance, security, encryption, and compliance requirements at scale.
- Relevant certifications across PostgreSQL, MongoDB, Databricks, or major cloud providers.
- Prior experience mentoring engineers or leading platform initiatives.
Key Skills & Attributes
- Technically savvy — genuinely strong and current across all three platforms, with the instincts to diagnose and optimize complex systems.
- Highly collaborative — works closely with engineering, analytics, and product teams and communicates clearly with both technical and non-technical audiences.
- Eager to learn and grow — curious, adaptable, and motivated to expand skills as the data estate and technologies evolve.
- Proactive and ownership-minded — anticipates problems, prevents outages, and takes end-to-end responsibility for platform health.
- Detail-oriented and reliable — disciplined about maintenance, documentation, and operational rigor.