Our crawlers collect regulatory updates from 80+ sources and push them through an automated
pipeline: extraction, embedding, search, classification, consolidation and translation. As we add
regions, keeping it healthy has become a job of its own.
We want an engineer who makes production tell us what's wrong before customers do, fixes what
can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach
a developer, it should arrive with a root cause and a suggested fix.
Not a ticket-driven ops role: you'll write production Python from week one and own platform
reliability alongside a core-team developer
Reliability & Observability Engineer (m/f/d)
München - HQ
Full-time
Permanent employee
75,000 - 90,000 € per year
About the Job
What you will do
• Cut the noise. Separate transient failures (network blips, timeouts, errors that vanish on rerun) from real ones, and build alerting the team trusts.
• Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right
developer.
• Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. "every source checked on time", "every
document produced text and embeddings".
• Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, deadletter handling, clear escalation when automation gives up.
• Keep agents safe: scoped permissions, audit trails, human approval for risky actions, and
measuring how often they're right.
• Continuously audit our Infrastructure and Identify opportunities to make it more efficient and save costs.
• Catch memory, timeout and cost problems across Cloud Run before they become silent
OOM kills.
• Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.
• Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right
developer.
• Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. "every source checked on time", "every
document produced text and embeddings".
• Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, deadletter handling, clear escalation when automation gives up.
• Keep agents safe: scoped permissions, audit trails, human approval for risky actions, and
measuring how often they're right.
• Continuously audit our Infrastructure and Identify opportunities to make it more efficient and save costs.
• Catch memory, timeout and cost problems across Cloud Run before they become silent
OOM kills.
• Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.
Our Tech Stack
Python, Docker, GCP (Cloud Run, Cloud Logging, Cloud Monitoring), Sentry, MongoDB, Azure
Blob, LLM APIs. Some Go and React.
Blob, LLM APIs. Some Go and React.
Your profile
• 5+ years in software engineering, SRE or platform roles, with real production Python.
• A track record of turning an ignored alert channel into one people act on.
• Hands-on experience building with LLMs or agents in production, and judgment about when
a plain if-statement is better.
• Solid GCP (or similar), containers and CI/CD.
• Strong grasp of distributed-system failure modes: retries, idempotency, partial failure,
backpressure.
• Pragmatism, and clear communication in a small team.
Nice to have
OpenTelemetry, SLOs, Terraform or Pulumi, data-quality observability, crawling or document processing.
• A track record of turning an ignored alert channel into one people act on.
• Hands-on experience building with LLMs or agents in production, and judgment about when
a plain if-statement is better.
• Solid GCP (or similar), containers and CI/CD.
• Strong grasp of distributed-system failure modes: retries, idempotency, partial failure,
backpressure.
• Pragmatism, and clear communication in a small team.
Nice to have
OpenTelemetry, SLOs, Terraform or Pulumi, data-quality observability, crawling or document processing.
Why us?
- Hybrid work culture: join us in our Munich office (min. 2 days/week).
- Flexible working hours.
- 26+4 vacation days per year (4 fixed “company rest days” over Christmas).
- 30 days of “workation” per year, within the EU and selected countries.
- High autonomy and flat hierarchies.
- EGYM Wellpass for unlimited access to fitness courses and gyms.
- Udemy access for educational videos.
Closing
We welcome candidates from all backgrounds and encourage diversity in our team. We encourage female and diverse engineers to apply and join our mission-driven culture that values open communication, work-life balance, and a welcoming environment. Convince us with your personality and your skills, and together we will make great things happen!
About us
Certivity is a Munich-based RegTech startup using AI to turn global regulatory complexity into clear, structured, and actionable compliance intelligence. Our SaaS platform automatically gathers and analyzes regulatory data worldwide, enabling engineering and compliance teams to work faster and more confidently.
Our international team of 100+ experts supports over 15,000 users in 10 countries. While we are strongly established in the automotive sector, we are now expanding into new industries with high regulatory demands.
Our international team of 100+ experts supports over 15,000 users in 10 countries. While we are strongly established in the automotive sector, we are now expanding into new industries with high regulatory demands.
