Staff AI Platform Engineer – Agent & Retrieval Infrastructure
Posted 3ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff AI platform engineer building secure Amazon Bedrock agent and retrieval infrastructure. Enabling Bedrock Ocean’s autonomous ocean-data platform and customer intelligence products.
Responsibilities:
- Design the Amazon Bedrock integration, including agent and action group configuration, backend APIs, model access, throughput, and cross-environment deployment
- Own the retrieval pipeline from ingestion and chunking through embedding and storage in Amazon OpenSearch Serverless
- Optimize index design, cost, and capacity
- Adapt ingestion pipelines for internal knowledge, ocean data, and customer platforms
- Address geospatial and large-binary dataset challenges
- Implement security with Bedrock Guardrails, VPC and PrivateLink boundaries, least-privilege IAM, and audit trails
- Define governance mechanisms for approval boundaries and safe autonomous agent actions
- Implement LLMOps and observability for tracing, tool calls, and retrieval performance using CloudWatch and tools such as Langfuse or Phoenix
- Build automated evaluation infrastructure, track results, and manage model-accuracy release gates
- Provide abstraction layers, SDKs, and self-service environments for independent AI feature delivery
- Manage infrastructure as code across environments
- Ensure deployment safety and participate in incident reviews
- Define the sequencing of internal, operational, ocean-data, and customer-facing retrieval capabilities
- Collaborate with the core data transport team on retrieval and data-freshness requirements
Requirements:
- 8+ years in software and infrastructure engineering
- Deep production backend experience; Python or TypeScript preferred, Go acceptable
- Staff-level ownership of technical direction
- Hands-on experience standing up Amazon Bedrock in production, including agents, knowledge bases, guardrails, model access, throughput, and quotas
- Containerized service deployment on ECS, EKS, or Lambda
- Ownership of CI/CD
- Practical RAG and vector search experience, including embeddings, chunking strategies, semantic search quality, and managed vector databases at production scale and cost
- Experience building or substantially extending ingestion pipelines over messy, heterogeneous, unstructured sources
- Strong AWS expertise: IAM, VPC networking, PrivateLink, Lambda, S3, KMS, CloudWatch, and infrastructure as code using Terraform, CDK, or CloudFormation
- Production experience with LLM features or autonomous agents
- Experience securing agentic systems, including tool permissions, prompt injection, exfiltration risk, sensitive data handling, and human-in-the-loop controls
- Experience designing developer-facing APIs, SDKs, or platform services with an API-first mindset
- Experience building and operating multi-tenant services with customer data isolation
- Demonstrated staff-level technical leadership and system design judgment
- Pragmatic preference for managed infrastructure
- Ability to work across several functions on a small team and document systems thoroughly
- Legally authorized to work in the US
- Eligible to obtain a government security clearance if required
Benefits:
- Equity
- Equal opportunity employer


















