Platform Engineer

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Platform Engineer building Kubernetes-based, multi-cloud infrastructure for Ema’s enterprise Agentic AI platform. Owning reliability, observability, distributed systems, and platform services.

Responsibilities:

  • Design, own, and evolve scalable multi-tenant microservices architectures on Kubernetes across GCP, Azure, and AWS
  • Build core platform and data-plane components in Golang and Python for data ingestion, knowledge-base indexing, vector/graph search, application connectivity, workflow automation, and ML operations
  • Own service-to-service communication, including gRPC/protobuf contracts, service mesh, load balancing, retries, timeouts, and circuit breaking
  • Document architectural tradeoffs involving partitioning/sharding, consistency models, caching, and build-vs-buy decisions
  • Define reliability contracts, including SLIs/SLOs, error budgets, capacity planning, autoscaling, and graceful degradation
  • Design and operate observability using Prometheus, Grafana, OpenTelemetry, distributed tracing, and real-time alerting
  • Drive DevOps and platform-engineering practices using Terraform, Helm, GitOps, and CI/CD pipelines
  • Optimize performance and cost through profiling, load testing, latency budgets, and cost-per-request analysis
  • Participate in on-call rotations and lead incident response and root-cause analysis

Requirements:

  • Bachelor's degree in Computer Science or a related field
  • 5+ years of experience in Platform, Infrastructure, or Backend Engineering
  • Strong CS fundamentals: data structures, algorithms, operating systems, and networking
  • Proficiency in Golang and Python
  • Production experience with Docker, Kubernetes, and microservices architecture
  • Hands-on experience with at least one major cloud provider (GCP, Azure, or AWS); multi-cloud a strong plus
  • Strong database expertise: query and read/write-path optimization, partitioning/sharding, replication and consistency models, with practical experience in NoSQL and graph stores
  • Solid grasp of the CAP theorem and database internals
  • Solid distributed-systems foundation: idempotency, backpressure, delivery semantics, and message queues such as Kafka, Pulsar, NATS, or PubSub
  • Track record of building platforms from the ground up that other engineering teams successfully build on
  • Experience operating systems at high scale
  • Depth in auth and security, including Vault, mTLS, RBAC, OIDC/SAML, and network policy
  • Experience with vector databases such as pgvector, Pinecone, or Milvus and graph databases such as Neo4j or Neptune
  • Open-source contributions to infrastructure projects, such as Kubernetes operators

Benefits:

  • Certain roles may be eligible for variable compensation, equity, and benefits
  • Equal employment opportunity commitment