You are one step closer
Interested in joining Langoor? Apply here and get an excuse to do monkey business.
Allowed file types: pdf, doc, docx, txt, text, rtf
Platform / DevOps Engineer, Client Embedded Pod (Kubernetes / Observability / Data Infrastructure)
Experience: 5+ years, Mid-Senior Level Positions / Openings: 5
About the role
You will join an embedded pod building the platform foundations that product and engineering teams run on, directly inside a client environment on live production systems. This is not a maintenance role. In a single quarter you could be hardening Kubernetes cluster lifecycle and capacity, tuning ClickHouse for high-cardinality observability, and standing up or scaling Milvus for a retrieval or agent workload. The pod moves fast and the client expects production-grade infrastructure, not one-off scripts.
What we expect from you
- You write clean, production-ready infrastructure and platform code, and you ship it fast. You do not need heavy process to move.
- You pick up unfamiliar stacks and tools quickly and without much hand holding – Kubernetes operators, GitOps, telemetry pipelines, columnar and vector databases.
- You already use modern AI tools to design, debug, and operate systems, and can show what that has done to your output.
- You are comfortable with ambiguity and switch context across infrastructure domains – cluster work one week, data-store tuning the next – without dropping reliability.
- You can hold a technical conversation directly with a client stakeholder, not just with your own team.
- You consistently turn platform work into measurable outcomes: developer velocity, cost control, and uptime.
Must-have skills
- Kubernetes in production: cluster lifecycle management, Helm charts and operators, capacity planning, and multi-tenant platform design
- GitOps workflows for infrastructure and application delivery (ArgoCD, Flux, or similar)
- ClickHouse or a comparable columnar store: schema design, ingestion pipelines, and query performance tuning for high-cardinality observability data (logs, traces, metrics, wide events)
- Vector databases in production – Milvus or similar (Pinecone, Weaviate, pgvector) – operating data planes for retrieval and agent workloads
- Telemetry pipeline design spanning logs, traces, and metrics
- Demonstrated use of AI tools in your engineering workflow, and awareness of how to apply AI capabilities to platform operations
- Git and version control workflows in a team setting (branching, PRs, code review)
Good-to-have skills
- Writing or extending Kubernetes operators and custom controllers
- Experience with complementary observability tooling (Prometheus, Grafana, OpenTelemetry)
- Cloud platform experience – AWS, Azure, or GCP
- CI/CD pipeline ownership, not just exposure
- Cost optimization or capacity forecasting for infrastructure spend
- Exposure to building AI agents or working with LLM APIs
- Comfortable writing infrastructure tests and validation as part of your normal workflow
- Experience working in fast, iterative cycles (agile or otherwise) with shifting priorities
- Prior experience working directly embedded with a client rather than purely internal teams
What this role is not
- This is not a role for someone who prefers a fixed scope and stable domain. The problem domain changes – cluster hardening one week, ClickHouse tuning the next, Milvus scaling after that.
- This is not an entry-level supervised role. You will be expected to operate independently within a few weeks of starting.
