Senior Elasticsearch Engineer
Posted 6hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Elasticsearch Engineer managing Elasticsearch and OpenSearch infrastructure at Chess.com. Owning the full data platform lifecycle with a focus on performance, reliability, and cross-team enablement.
Responsibilities:
- Own the full lifecycle of our search and analytics data platform: capacity planning, cluster architecture, performance tuning, incident response, migration strategy, and operational excellence.
- Be the single point of deep expertise across all Elasticsearch and OpenSearch clusters at Chess.com.
- Hands-on with cluster internals, write ILM/ISM policies, push infrastructure changes through GitOps, and make real-time decisions about replica allocation when a cluster goes red.
- On-call ownership for Elasticsearch-related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades.
- Real-time cluster triage and cross-team coordination during production incidents.
- Post-mortem authoring and systemic reliability improvements.
- Snapshot and disaster recovery management across clusters.
- Advise engineering teams on index design, mapping strategy, retention policies, and query optimization.
- Manage Kibana and OpenSearch Dashboards access and configuration for internal consumers.
- Define and maintain workload priority tiers across clusters.
Requirements:
- 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput)
- Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management
- Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments
- Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management)
- Hands-on Linux systems administration with a focus on storage and I/O performance
- Experience managing both Elasticsearch and OpenSearch in production, including an informed opinion on their respective trade-offs
- Incident command experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating clearly to stakeholders
- Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code for cluster configuration
- Fluency with the Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore.
Benefits:
- Full autonomy.
- Real scale.
- Bare metal.
- Strategic impact.
- Small team, big trust.

















