Senior Elasticsearch Engineer
Chess.com
- DevOps-Engineer
- Developer
- Elasticsearch-Data-Engineer
- Elasticsearch-Engineer
- Elasticsearch-Specialist
- Infrastructure-Engineering
- Search-Engineering
- Senior-Search-Engineer
- Site-Reliability-Engineering
Assessed from original listing evidence
The role
Job description
About Us
[link removed] is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying chess.
We are a team of 600+ fully remote people in 60+ countries working hard to serve the global chess community. We are here to support 250M+ chess players worldwide with the best possible product, content, and tools to serve the community!
We are a tech company. A gaming company. A content company. And we do it all with passion and commitment to the game. Above all we prize our mission-driven, flat, life-celebrating, no-corporate culture, and we look forward to meeting you and learning more about what you can bring to the team.
About You
[link removed] is the world's largest chess platform with 235M+ members and ~20 million daily games. Our Elasticsearch and OpenSearch infrastructure underpins search, user activity, analytics, logging, and operational intelligence at massive scale, hundreds of terabytes across a dozen production clusters running on bare-metal Kubernetes.
We're looking for a Senior Elasticsearch Engineer who can own the full lifecycle of our search and analytics data platform: capacity planning, cluster architecture, performance tuning, incident response, migration strategy, and operational excellence. You'll be the single point of deep expertise across all Elasticsearch and OpenSearch clusters at [link removed].
This is not a monitoring-from-dashboards role. You'll be hands-on with cluster internals, write ILM/ISM policies, push infrastructure changes through GitOps, and make real-time decisions about replica allocation when a cluster goes red.
What you'll do
Incident Response & Reliability
- Shard allocation strategy for write-heavy data streams at high throughput (millions of documents per minute)
- Disk watermark management, retention policy tuning, and rollover orchestration for high-volume indices
- Performance optimization and I/O tuning on bare-metal nodes
- Write queue analysis, thread pool diagnostics, and shard rebalancing under load
- Capacity planning and growth forecasting across clusters
Incident Response & Reliability
- On-call ownership for Elasticsearch-related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades
- Real-time cluster triage and cross-team coordination during production incidents
- Post-mortem authoring and systemic reliability improvements
- Snapshot and disaster recovery management across clusters
Migration & Strategy
- Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation across ILM/ISM, security models, and plugin ecosystems
- Version upgrade planning and rolling restart orchestration with zero-downtime requirements
- End-to-end new cluster provisioning and onboarding
Cross-Team Enablement
- Advise engineering teams on index design, mapping strategy, retention policies, and query optimization
- Manage Kibana and OpenSearch Dashboards access and configuration for internal consumers
- Define and maintain workload priority tiers across clusters
Preferred Skills
- 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput)
- Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management
- Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments
- Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management)
- Hands-on Linux systems administration with a focus on storage and I/O performance
- Experience managing both Elasticsearch and OpenSearch in production, including an informed opinion on their respective trade-offs
- Incident command experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating clearly to stakeholders
- Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code for cluster configuration
- Fluency with the Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore
Bonus Experience
- OpenSearch ISM policies and security plugin (fine-grained access control)
- GCS or S3 snapshot repository configuration and cross-cluster replication
- Grafana + Prometheus monitoring for Elasticsearch metrics
- Kibana Discover, Dev Tools, and data view management at scale
- Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers)
- Vault integration for secrets management in Kubernetes-deployed search clusters
- Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch
- Hardware selection experience for search-optimized server configurations
- Python or scripting for operational analysis and automation
What Makes This Role Special
- Full autonomy. You are the Elasticsearch authority. You make the architecture calls, set the priorities, and own the outcomes.
- Real scale. Hundreds of terabytes of data, billions of documents, millions of daily queries. The problems here don't exist at smaller companies.
- Bare metal. No managed Elastic Cloud. You're operating directly on the hardware. This is hands-on engineering.
- Strategic impact. Your decisions on ES vs. OpenSearch migration, cluster topology, and capacity planning directly affect product capabilities and infrastructure costs.
- Small team, big trust.[link removed] runs lean. You won't be buried in process or approvals. Ship changes, fix problems, improve systems.
About the Opportunity
- This is a full-time opportunity
- We are 100% remote (work from anywhere!)
---
You can learn more about us here:
- [link removed]
- [link removed]
Originally posted on Himalayas
Keep exploring
Related remote jobs
Database Engineer
Chess.com
Data EngineeringAndroid Engineer
Chess.com
Software EngineeringSenior Product Manager
Chess.com
Product ManagementSenior Conversion Copywriter
Chess.com
Writing & ContentSenior Product Manager, Chess AI
Chess.com
Machine Learning