Skip To Main Content
backBack to Search

Senior Platform Engineer (AI Platforms)

Remote in Brazil, & 4 others
Platform Engineering& 8 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We are looking for a Senior Platform Engineer to build shared AI platform services that power LLM integrations and agent runtimes, with observability, guardrails, and cost attribution designed in from day one. You will run and evolve Kubernetes-based gateway, MCP, and telemetry capabilities that application teams rely on.

Responsibilities
  • Build and operate shared AI platform services including LLM gateways, routing, fallback, MCP layers, and agent runtimes
  • Design and run model onboarding across environments with routing rules, quotas, tiering, and version deprecation paths
  • Implement end-to-end observability for platform components using OpenTelemetry, metrics, and structured logs
  • Implement and maintain Langfuse telemetry including traces, spans, prompts, completions, feedback, and evaluation metadata
  • Design and operate the MCP layer for tool registration, discovery, session handling, permission boundaries, and tool-call tracing
  • Run platform services on Kubernetes with GitOps using Helm and Kustomize, autoscaling, health checks, and progressive rollout
  • Implement authentication, authorization, and tenancy with OAuth2/OIDC, SSO, JWT validation, and API key management
  • Apply guardrails and policy enforcement including PII detection, filtering, audit logging, and retention rules
  • Build dashboards, alerts, and reports for reliability, latency, errors, quality signals, and production behavior
  • Design cost and usage observability across model providers with attribution by app, team, user, model, and environment
  • Create showback-ready metrics for token usage, model mix, cache behavior, and vendor spend to guide routing and capacity
  • Support agent lifecycle capabilities including publishing, versioning, discovery, memory, scaling, and resilience
  • Build developer self-service workflows for onboarding, provisioning, templates, and internal portals or CLIs
Requirements
  • Senior-level platform engineering experience (3+ years) with AI platform services
  • Strong Kubernetes experience (3+ years) operating production services
  • Hands-on Langfuse experience implementing AI observability patterns
  • Strong Model Context Protocol (MCP) experience building or operating MCP servers and registries
  • Strong OpenTelemetry experience for distributed tracing, metrics, and structured logs
  • Production GitOps experience using Helm and Kustomize for declarative delivery
  • Strong IAM and security knowledge covering OAuth2/OIDC, SSO, JWT, and secrets handling
  • Strong communication skills to partner with engineering and developer teams
  • Upper-Intermediate English proficiency (B2) for collaboration and documentation
Nice to have
  • Hands-on LangChain experience for agent or tool integration patterns
  • Hands-on LangGraph experience for graph-based agent workflows
  • Retrieval-Augmented Generation (RAG) experience with evaluation and retrieval quality tuning