Platform Engineer
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Platform Engineer building Kubernetes-based, multi-cloud infrastructure for Ema’s enterprise Agentic AI platform. Owning reliability, observability, distributed systems, and platform services.
Responsibilities:
- Design, own, and evolve scalable multi-tenant microservices architectures on Kubernetes across GCP, Azure, and AWS
- Build core platform and data-plane components in Golang and Python for data ingestion, knowledge-base indexing, vector/graph search, application connectivity, workflow automation, and ML operations
- Own service-to-service communication, including gRPC/protobuf contracts, service mesh, load balancing, retries, timeouts, and circuit breaking
- Document architectural tradeoffs involving partitioning/sharding, consistency models, caching, and build-vs-buy decisions
- Define reliability contracts, including SLIs/SLOs, error budgets, capacity planning, autoscaling, and graceful degradation
- Design and operate observability using Prometheus, Grafana, OpenTelemetry, distributed tracing, and real-time alerting
- Drive DevOps and platform-engineering practices using Terraform, Helm, GitOps, and CI/CD pipelines
- Optimize performance and cost through profiling, load testing, latency budgets, and cost-per-request analysis
- Participate in on-call rotations and lead incident response and root-cause analysis
Requirements:
- Bachelor's degree in Computer Science or a related field
- 5+ years of experience in Platform, Infrastructure, or Backend Engineering
- Strong CS fundamentals: data structures, algorithms, operating systems, and networking
- Proficiency in Golang and Python
- Production experience with Docker, Kubernetes, and microservices architecture
- Hands-on experience with at least one major cloud provider (GCP, Azure, or AWS); multi-cloud a strong plus
- Strong database expertise: query and read/write-path optimization, partitioning/sharding, replication and consistency models, with practical experience in NoSQL and graph stores
- Solid grasp of the CAP theorem and database internals
- Solid distributed-systems foundation: idempotency, backpressure, delivery semantics, and message queues such as Kafka, Pulsar, NATS, or PubSub
- Track record of building platforms from the ground up that other engineering teams successfully build on
- Experience operating systems at high scale
- Depth in auth and security, including Vault, mTLS, RBAC, OIDC/SAML, and network policy
- Experience with vector databases such as pgvector, Pinecone, or Milvus and graph databases such as Neo4j or Neptune
- Open-source contributions to infrastructure projects, such as Kubernetes operators
Benefits:
- Certain roles may be eligible for variable compensation, equity, and benefits
- Equal employment opportunity commitment
















