Lead Software Engineer - Agent Safety

Kraken

Berlin, DEhybridPosted Aug 25, 2026
Posting intelligenceActively listed

Skills

langchaindatadogdjangopythonopenaiawsllm

About the role

Help us use technology to make a big green dent in the universe!

Kraken powers some of the most innovative global developments in energy.

We create the technology that redefines utilities and unlocks a new energy system of the future. By optimising renewable generation, building a more intelligent grid, and empowering utilities to deliver an exceptional customer experience, our operating system is transforming the industry worldwide.

It’s an incredibly exciting time to work in energy. Join us on our mission to improve the lives of ONE BILLION humans within the decade and shape a cleaner, better future for everyone.

AI is a key investment area for Kraken Technologies as we look to our existing capabilities. A crucial part of this is broadening the foundational infrastructure to enable teams across the organisation to use AI effectively to accelerate our mission.

You’ll work in the Agent Safety/Evals team, a new sub-team within AI Foundations. We build the shared platforms, harnesses, and guardrails that enable engineering and product teams to safely, reliably, and deterministically use machine learning and generative AI agents for internal systems and workflows across the business. This is a delivery-focused team that sits at the intersection of engineering and delivery, focusing on empowering our internal users.

The Role

We’re hiring a Lead Software Engineer to head our newly formed Agent Safety/Evals team. As AI agents take on more autonomous tasks across Kraken's internal workflows, this role is critical to ensuring they do so securely and predictably.

This is a leadership role focused on constraining our internally facing AI agents to an expected operation space and guaranteeing system reliability.

You will define the technical strategy for how we evaluate models, enforce safety guardrails, and govern AI behavior across internal platforms, skills, and harnesses. You’ll work closely with the broader AI Foundations team and engineers across Kraken to ensure that our push for rapid AI adoption in internal tooling never compromises on security, determinism, or quality.

What you'll own

Lead the technical direction for internal agent safety: Design and implement systems focused on the reliability, security, and determinism of LLMs and autonomous agents, constraining them strictly to expected operational spaces within our internal ecosystems.

Build robust evaluation frameworks: Develop scalable harnesses and evals tailored to internal workflows and skills. You will be responsible for asking the right types of questions about the quality and reproducibility of our evals, while engineering the systems to measure them robustly.

Implement guardrails and governance: Create and enforce pre- and post-generation guardrails, managing the overarching governance of AI models operating within internal tools and platforms.

Drive AI security and observability: Build out dedicated auth/permissions for internal AI agents, establish deep monitoring/observability pipelines, and define incident response protocols for AI-specific anomalies.

Verification and Red Teaming: Lead continuous verification efforts and red teaming exercises to proactively identify vulnerabilities, prompt injections, or unpredictable behaviours in our internal AI implementations.

Operate in AWS: Deploy, run, and support high-throughput, low-latency safety services for internal use cases; make sensible architecture/cost tradeoffs; partner effectively with platform/techops/security stakeholders.

What you bring to the party

Strong technical leadership: Proven experience leading technical initiatives or teams, capable of setting the technical vision for complex, ambiguous domains.

Deep software engineering fundamentals: Senior/advanced capability in designing secure components end-to-end, testing thoroughly, and reasoning heavily about system design, concurrency, and architecture tradeoffs. (Python preferred).

Expertise in AI Evaluation and Safety: A highly critical thinker who understands the nuances of LLM behavior. You must be comfortable interrogating the quality of AI outputs and deeply experienced in building harnesses to measure it reproducibly.

Security and Governance mindset: Experience with threat modeling, authentication, red teaming, or building guardrails for internal systems, tooling, or platforms.

Cloud experience (AWS): Comfortable running internal services, owning reliability/scalability, and collaborating with platform/techops/security partners.

Excellent communication: Able to synthesize complex safety and eval metrics into actionable insights for both technical and non-technical stakeholders.

What Success Looks Like

Confidence in Internal Tooling: Engineers across Kraken can deploy new autonomous models, AI skills, and internal agents with total confidence, knowing they are bound by reliable safety guardrails and strict operational constraints within our internal ecosystems.

High-Quality Internal Evals: You have established a culture and infrastructure of robust, highly reproducible evaluation metrics that accurately reflect the real-world performance and safety of our internal-facing AI tooling and harnesses.

Resilient Internal Infrastructure: AI security, monitoring, and incident response are treated as first-class citizens, ensuring that any erratic agent behavior in our internal platforms is immediately caught and mitigated before it affects broader operations.

Technical Leadership: Strong collaboration across AI Foundations helps accelerate secure internal AI adoption; you are actively mentoring engineers and shaping the company's internal AI governance strategy.

️ Bonus points

Experience with specific LLM evaluation and safety frameworks (e.g., Inspect AI, Ragas, OpenAI Evals, NeMo Guardrails).

Django experience and strong backend engineering patterns (security, performance, maintainability).

Experience with Datadog for complex observability, tracing, and monitoring in AI environments.

Familiarity with foundational AI engineering tooling like Pydantic AI, LiteLLM, or LangChain.

Are you ready for a career with us? We want to ensure you have the right tools and environment to unleash your potential. If you have any accommodation requests, please contact us at inclusion@kraken.tech. We’ll do our best to tailor the experience to your needs so you can feel comfortable, confident and be at your best!

Our (i) Applicant and Candidate Privacy Notice and Artificial Intelligence (AI) Notice, (ii) Website Privacy Notice and (iii) Cookie Notice govern the collection and use of your personal data in connection with your application and use of our website. These policies explain how we handle your data and outline your rights under applicable laws, including, but not limited to, the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA). Depending on your location, you may have the right to access, correct, or delete your information, object to processing, or withdraw consent. By applying, you acknowledge that you’ve read, understood and consent to these terms

Please note that, in line with our current recruitment policy, we are unable to offer visa sponsorship for this position; applicants must have the right to work in the country that they're applying to, at the time of application.*

Questions about this role

Click "Apply with AI Applyd" above and you are done. Your resume is rewritten for this advert, the screening questions are answered, and it is submitted on Kraken's own hiring system. No retyping your history, no fourteen tabs, no evening lost.

Compensation for Software Engineer roles in Germany varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Software Engineer hub for Germany medians across recent openings.

You never touch the form - the application is filled and submitted for you on Kraken's own hiring system. It is not marked sent when we press submit. It is marked sent when a confirmation from their system arrives at the address we apply with, and your dashboard shows which stage each application is at until then.

Twelve applicant tracking systems have a real apply path: Workday, Greenhouse, Lever, Ashby, Workable, iCIMS, Personio, Recruitee, Teamtailor, Rippling, Breezy and SmartRecruiters. Your application goes in on the employer's own hiring system, never into an aggregator queue.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.