Oliver Dippel Agentic AI · Reinforcement Learning · Foundation Models

About

I am an AI researcher and engineer with a Ph.D. in Artificial Intelligence from the University of Liverpool. My work spans reinforcement learning, agentic AI, foundation models, and reliable machine-learning systems.

My doctoral research focused on deep reinforcement learning for continuous industrial processes, particularly under limited data, incomplete system knowledge, and safety constraints. My research includes transformer-based agents for goal-oriented and in-context reinforcement learning.

Alongside research, I have worked on production-facing AI systems. At ZeroAI, I developed and evaluated agentic LLM systems for safety-critical autonomous-driving workflows, including tool use, structured-output validation, adversarial evaluation, and failure containment.

I currently contribute to Birch, where I work on AI-powered product search and product question answering for e-commerce. This includes retrieval-augmented generation, evaluation and verification pipelines, multi-tenant system design, and explainable product ranking.

Earlier in my career, I worked in Early Development Biostatistics at Novartis, developing internal software to support scientific workflows. I hold an M.Sc. in Applied Statistics and a B.Sc. in Economics from the Georg-August University of Göttingen .

Beyond Work

Outside of work, I enjoy endurance sports, kitesurfing, snowboarding, and travelling. I approach both research and engineering with the same persistence, curiosity, and willingness to keep improving.

What I Work On

Agentic AI Systems Tool-using agents, workflow orchestration, structured outputs, failure containment, evaluation, observability, and reliable deployment.
Reinforcement Learning Sequential decision-making, continuous control, goal-conditioned learning, in-context reinforcement learning, and learning under uncertainty.
Foundation Models and RAG Retrieval systems, grounded question answering, semantic product search, verification pipelines, and hallucination-resistant applications.
Scalable Machine Learning PyTorch, JAX, Ray, distributed experimentation, reproducible pipelines, cloud deployment, and production monitoring.

Research Interests

My research interests are about developing autonomous learning systems that operate robustly in complex and uncertain environments. I am interested in agents that go beyond solving a single predefined task, and instead learn how to reason, adapt, and improve over time through interaction.

At the core of my interests lies the intersection of reinforcement learning, sequential decision-making, and large-scale models. I study how artificial agents can understand their surroundings, plan over extended time horizons, and generalize their behavior beyond the data or situations they were explicitly trained on.

A long-term motivation is the pursuit of artificial general intelligence (AGI)—systems capable of flexible reasoning, continual learning, and adaptation across diverse domains without retraining from scratch.

Ultimately, my goal is to bridge the gap between specialized AI systems and genuinely open-ended learners that can address a wide range of real-world challenges in a scalable and reliable manner.

Publications

Heuristic Transformer: Belief Augmented In-Context Reinforcement Learning Oliver Dippel, Alexei Lisitsa, Bei Peng · 2025
Contextual Transformers for Goal-Oriented Reinforcement Learning Oliver Dippel, Alexei Lisitsa, Bei Peng · 2024
Deep Reinforcement Learning for Continuous Control of Material Thickness Oliver Dippel, Alexei Lisitsa, Bei Peng · 2023
Note on Code Availability My Ph.D. research is partially funded by a corporate partner. Consequently, the production-grade codebase and proprietary libraries used for these publications are subject to Intellectual Property (IP) rights and are not publicly available on GitHub. Please feel free to reach out for a technical discussion regarding the underlying architectures and methodologies.