Benoit Seguin

Benoit Seguin

Staff Software Engineer at Google
Tokyo, Japan (relocating to Zurich in December)

I am a Staff Software Engineer at Google, where I co-lead Simula, Google's core synthetic data framework. Recently, I’ve focused on building agentic, self-improving pipelines for data and RL environment synthesis. I specialize in conceptualizing and engineering the composable abstractions that turn frontier research ideas into robust, high-performance systems at scale.


Featured Systems & Projects

Simula (2024–Present)

Co-Founder & Co-Lead | Google’s Core Synthetic Data Framework

Teaching AI models requires massive amounts of high-quality data, but generating it manually is impossible.

Simple automated prompting causes AI to create repetitive, low-quality data. I co-founded Simula with Hamza Harkous to solve this. We designed a programmable framework that independently controls data diversity, complexity, and quality without human intervention.

Our team scaled Simula into Google's primary internal multimodal data synthesis engine. It is now used by over 2,500 Googlers, was used to generate trillions of tokens, and achieved a 93% user satisfaction score, the highest among all data tools at Google. Used by hundreds of teams, today it is a key enabler for the Gemma ecosystem, provides the primary synthetic data backbone for Gemini safety classifiers, powers production user protection features like AI-powered scam detection, and extends much further across frontier AI applications and data use cases.

Highlights & Applications

Internal LLM Prompting & Agentic Bulk Inference Engine (2023–Present)

Creator & Lead Architect | Prompt Templating & Distributed Batch Inference

As LLM workflows evolved beyond simple chat prompts into complex agentic graphs and bulk synthetic generation, research teams needed a flexible way to manage prompt composition and run large-scale inference efficiently.

I created a lightweight, expressive prompt management toolkit and bulk inference engine that became popular among researchers across Google for rapid prototyping and large-scale batch generation.

Architecture & Highlights

  • Expressive Prompt Templating: Extends modern templating engines with composable primitives for structured schema enforcement, multimodal context assembly, dynamic few-shot injection, and modular partials.
  • Distributed Bulk Inference: Built horizontally scalable inference pipelines on top of Jax-on-beam, allowing complex multi-step and agentic generation workflows to execute across distributed accelerator clusters.
  • Off-Peak Compute Efficiency: Leveraged opportunistic, off-peak datacenter capacity to run high-throughput batch inference workloads with significant infrastructure savings.
  • Foundational Layer for Simula: Provided the core prompt rendering and batch inference substrate underlying Simula's data generation pipelines.

Scale & Impact

  • Reached 10,000+ monthly active users, becoming a go-to toolkit across Google for structured prompt templating and bulk inference.
  • Used to generate multi-trillion tokens across diverse research workflows and data generation pipelines.

Entrepreneurship & Cultural Heritage AI (2014–2022)

Before joining Google, my research and engineering focused on building large-scale computer vision systems for complex, real-world data collections, spanning startup entrepreneurship, academic research, and world-class cultural institutions.

ArtBeat.ai (2021–2022)

CTO & Co-Founder | AI-Powered Art Market Intelligence

  • Co-founded and led the engineering team (3 engineers) to construct an end-to-end intelligence and valuation engine for the global art market.
  • Data Ingestion Pipeline: Architected a scalable data crawling and normalization pipeline collecting millions of historical auction transactions, catalog raisonnés, gallery records, and high-resolution artwork images.
  • Multimodal Valuation Models: Developed an explainable multimodal valuation model that integrated visual image embeddings, historical price dynamics, artist network graphs, and auction house metadata—outperforming expert human estimates 45% of the time.
  • Production Architecture: Built and deployed the full cloud infrastructure, microservices backend, and customer-facing exploration application.

High-Impact Cultural Heritage Consulting & Academic Leadership (2019–2022)

Independent ML Consultant & Lecturer at ETH Zurich

  • Getty Research Institute (Los Angeles): Formulated the technical strategy and computer vision architecture to automate the digitization, document layout analysis, and catalog organization of the Getty's massive Photo Archive.
  • ETH Library (Zurich): Designed and implemented a full-stack text-reuse and cross-correlation platform (e-rara) analyzing 500,000+ digitized pages of architectural history to track intellectual influence across centuries.
  • The Impresso Project (EPFL & Luxembourg): Built a high-performance visual search engine and recommendation system indexing millions of historical newspaper photographs and media archives.
  • Lecturer at ETH Zurich: Taught advanced applied machine learning and data science courses.

EPFL PhD: The Replica Project & dhSegment (2014–2018)

PhD Researcher | Digital Humanities Laboratory (DHLAB), EPFL
Advisor: Prof. Frédéric Kaplan

  • The Replica Project: Spearheaded the computer vision architecture for a flagship initiative with the Cini Foundation in Venice, developing deep-learning visual similarity algorithms to discover composition borrowing, student copies, and visual citations across hundreds of thousands of Renaissance artworks.
  • dhSegment: Co-created and open-sourced dhSegment, a versatile pixel-level deep learning framework for historical document segmentation and layout analysis, widely adopted across digital humanities laboratories globally.
  • Best Demonstration Award: Research Days of the Computer Science Faculty, EPFL (2017).

Career Experience

  • Google | Staff Software Engineer (2025 – Present), Senior Software Engineer (Sept 2022 – 2025)
    Co-founder & Co-lead of Simula; Creator of Google's internal LLM Prompting & Bulk Inference Engine; Architect of Leap AI; Core Tech Impact Award Winner; Top-2 Code Author (2024).
  • ArtBeat.ai | CTO & Co-Founder (Mar 2021 – Apr 2022)
    Led engineering for an AI-driven art market valuation startup; architected multimodal valuation models and multi-source ETL pipelines.
  • Benoit Seguin Consulting & Software Development | Principal Consultant (Jan 2019 – Apr 2022)
    Machine learning consulting for world-class cultural heritage institutions: Getty Research Institute, ETH Library, and the Impresso Project.
  • ETH Zurich | Lecturer (2019 – 2021)
    Taught practical machine learning and scalable data processing.
  • EPFL (DHLAB) | PhD Candidate & Researcher (Sept 2014 – Nov 2018)
    Conducted research on deep learning for large iconographic archives; authored Replica and dhSegment.

Education

  • Ph.D. in Computer ScienceEPFL (Swiss Federal Institute of Technology, Lausanne) (2014–2018)
    Thesis: Making large-scale art historical photo archives searchable: A deep learning approach.
  • M.Sc. in Computer ScienceEPFL (2011–2013)
    Specialization in Computer Vision, Machine Learning, and Distributed Systems.
  • Diplôme d’IngénieurÉcole Polytechnique (Paris) (2008–2012)
    France's premier engineering Grande École; intensive curriculum in Applied Mathematics and Computer Science.

Selected Publications & Technical Reports

View full publication list on Google Scholar

Reasoning-Driven Synthetic Data Generation and Evaluation
Transactions on Machine Learning Research (TMLR) / arXiv:2603.29791 (2026)
Davidson, T. R., Seguin, B., Bacis, E., Ilharco, C., Harkous, H.

Introduces Simula, reframing synthetic data generation as structured mechanism design and first-principles reasoning rather than heuristic prompt hacking, providing programmable control over diversity, complexity, and dataset fidelity.

Gemma 4 Technical Report
Google DeepMind (2026)
Gemma Team (incl. Seguin, B.)

Technical report detailing Google's next-generation open-weights model family. Contributor to the high-throughput synthetic data generation pipelines powered by Simula.

dhSegment: A Generic Deep-Learning Approach for Document Segmentation
16th International Conference on Frontiers in Handwriting Recognition (ICFHR) (2018)
Oliveira, S.*, Seguin, B.*, Kaplan, F.

Presents a general-purpose, convolutional neural network architecture for historical document layout analysis, page extraction, and text line segmentation.

The Replica Project: Building a Visual Search Engine for Art Historians
XRDS: Crossroads, The ACM Magazine for Students (2018)
Seguin, B.

Invited overview of the computer vision architecture, reverse image indexing, and visual link retrieval mechanisms developed for the Cini Foundation's iconographic collections in Venice.

New Techniques for the Digitization of Art Historical Photographic Archives
Archiving Conference, IS&T (2018)
Seguin, B., Costiner, L., di Lenardo, I., Kaplan, F.

Describes the automatic image processing and computer vision pipeline deployed for the digitization and enrichment of the Giorgio Cini Foundation's photo collection in Venice.

Visual Link Retrieval in a Database of Paintings
ECCV Visart Workshop (2016)
Seguin, B., Striolo, C., di Lenardo, I., Kaplan, F.

Formulates the metric learning framework for visual similarity search and visual citation discovery across large collections of fine art.

Deep Learning for Logic Optimization Algorithms
IEEE International Symposium on Circuits and Systems (ISCAS) (2018)
Haaswijk, W.*, Collins, E.*, Seguin, B.*, Soeken, M., Süsstrunk, S., Kaplan, F., De Micheli, G.

Explores the application of deep reinforcement learning for logic synthesis and combinatorial Boolean network optimization.