Database Administrator (Bengaluru)

Database Administrator (Bengaluru)

29 Aug
|
SSGServ™
|
Bengaluru

29 Aug

SSGServ™

Bengaluru

DATABASE ADMINISTRATOR

Role :Database Administrator (MySQL / Percona)

Experience : 3+ Years (Immediate Joiner)

Shift :US Eastern Time business hours night shift — approximately 6:30 PM to 3:30 AM IST, shifting by about an hour with US daylight saving.

Budget : 6 LPA - 8LPA

Location : Banglore

About the roleClient runs MySQL / Percona replication clusters that serve production workloads for hundreds of real estate client sites. This is a hands-on engineering role — you own database problems from detection through resolution, not NOC-tier monitoring.

The Database Administrator is responsible for the health, performance, and recoverability of these clusters. The role demands deep MySQL expertise, the ability to diagnose and remediate replication failures, and a disciplined approach to backup verification and capacity management. Solid Linux administration is required, since all database workloads run on Linux VMs.

Workplace● Database: MySQL / Percona Server — multiple primary/replica clusters (moya cluster, listings cluster, separate topologies)

● OS / virtualization: Oracle Linux / AlmaLinux (Red Hat family) on XCP-ng VMs

● Backup: Percona XtraBackup

● Monitoring: Zabbix with custom MySQL templates

● Automation: Ansible, Bash/Python scripts

● Storage: LVM volumes on XCP-ng SR, Dell PowerStore iSCSI backend

Key responsibilities● Replication health — SSH to each replica daily; verify Slave_IO_Running and Slave_SQL_Running are Yes and Seconds_Behind_Master is at or near zero. When replication breaks, identify the error, remediate or skip the offending transaction safely, and document the incident.

● Replication diagnosis — Interpret Last_SQL_Error, classify error types (duplicate key 1062, row-not-found 1032, definition mismatch), inspect binlog position with mysqlbinlog, and judge when to skip an event versus rebuild the replica. Escalate data-loss-risk decisions.

● Performance tuning — Review the slow query log, use EXPLAIN to analyze execution plans, identify missing indexes and full table scans, implement index changes, and tune innodb_buffer_pool_size (target 70–80% on dedicated hosts).

● Backup & recovery — Run and verify Percona XtraBackup jobs (--backup, --prepare), confirm clean logs, perform periodic test restores to a non-production instance, and maintain a recovery runbook with an RTO estimate.

● Capacity planning — Track disk-usage trends and ibdata1 / binlog growth, project storage exhaustion,



and coordinate LVM expansion with the sysadmin before hitting 80% capacity.

● User & privilege management — Create MySQL users under least-privilege, audit with SHOW GRANTS, remove unused accounts, document service credentials in Bitwarden, and rotate passwords on schedule.

● Alert response — Respond to Zabbix alerts (replication lag, disk, connection count, slow-query rate, service down); confirm the condition and root-cause with SHOW STATUS, SHOW PROCESSLIST, and SHOW ENGINE INNODB STATUS.

● Schema & query support — Review proposed schema changes for performance and index coverage, use pt-online-schema-change for large alterations, test DDL on staging, and provide EXPLAIN analysis with written rationale.

● Binlog management — Maintain binlog retention (expire_logs_days / binlog_expire_logs_seconds) sufficient for replication tolerance and point-in-time recovery, and purge old binlogs safely once all replicas have consumed them.

● Linux administration — Perform disk/LVM expansion, package updates, service and cron management, and firewalld rules for port 3306 on the database VMs; review OS logs for issues affecting database stability.

● Incident response — Follow the runbook during incidents, capture status output immediately, check for disk-full / OOM conditions, remediate within SLA, and write a post-mortem (timeline, root cause, corrective and preventive action) in Wiki.js.

● Automation — Write and maintain cron-driven Bash/Python for replication health checks, backup verification, disk trending, and slow-query summaries, stored in Git with a README.

● Version upgrades — Plan and execute staged major upgrades along the required 5.6 → 5.7 → 8.0 → 8.4 path (never skipping a version), validating at each step. Handle known breaking changes — sql_mode, authentication plugin, reserved keywords, deprecated variables — upgrade replicas first, then fail over and upgrade the former primary, with a tested rollback plan.

You should be able to, without assistance● Connect to a replica, identify that replication is broken,



determine the error, and safely resume replication.

● Use EXPLAIN to analyze a slow query and propose a specific index to resolve a full table scan.

● Calculate an appropriate innodb_buffer_pool_size for a host with 32 GB RAM running only MySQL.

● Initiate and verify a Percona XtraBackup, then restore it to a test instance.

● Write a Bash script that checks SHOW SLAVE STATUS and emails an alert when replication lag exceeds a threshold.

● Expand a disk volume on a Linux VM using LVM without service interruption.

● Describe the MySQL 5.6-to-8.4 upgrade path, name at least three breaking changes between major versions, and explain the replica-first strategy.

Required skills● Deep MySQL / Percona expertise: replication, performance tuning, and recovery.

● Percona XtraBackup for backup and restore, with verified test restores.

● Query analysis with EXPLAIN, slow-query logs, and performance_schema.

● Solid Red Hat–family Linux administration, including LVM.

● Bash and Python scripting; Git-based workflow.

● Zabbix monitoring and alert response for database workloads.

Nice to have● pt-online-schema-change and the wider Percona Toolkit.

● Experience executing staged MySQL major-version upgrades in production.

● XCP-ng virtualization and Dell PowerStore / iSCSI storage.

Shared expectations● Communication — Provide daily status updates on open issues and flag blockers proactively, before a deadline is missed. Clear written English is required; technical shorthand is fine, ambiguity is not.

● Escalation — Any production-impacting issue not resolved within 30 minutes is escalated to the Lead Systems Engineer with a summary of what is broken, what has been tried, and current status.

● Change control — No production changes outside approved maintenance windows without Lead Engineer approval. Every change, including minor config edits, is documented in the team wiki before and after execution.

● Documentation — Runbooks, configs, and post-mortems must be written so another engineer of equivalent experience can follow them cold. Incomplete documentation is treated as incomplete work.

● Security hygiene — Credentials are never stored in scripts, config files, or wiki pages in plain text; all service-account passwords live in Bitwarden. Suspected security incidents are reported to the Lead Engineer immediately.

📌 Database Administrator (Bengaluru)
🏢 SSGServ™
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: database administrator (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: database administrator (bengaluru) / bengaluru