A wet lab for AI-ready spatial biology datasets in oncology

Explicyte generates production-scale spatial transcriptomics, single-cell, and digital pathology datasets from human tissues sourced through accredited biobanks.

We deliver in open formats, ready for ingestion into your AI pipeline.

Supporting biology foundation models

Dataset generation, secured end-to-end

Cohort Design & Sourcing

Unique cohorts and safe-to-use data, sourced for you

  • Cohort design with PhD-level expertise — indications, controls, stratification, sample size for statistical power
  • Multiple sample types — tumor and normal FFPE tissues, blood-derived samples
  • Sourcing through accredited French biobanks (NF S96-900, ISO 20387)
  • Anonymized metadata — clinical, histopathological, epidemiological — usable for AI training
About tissue sourcing

Sample QC & processing

Wet lab rigor to process the right samples & minimize batch effects

  • ISO-certified QMS — sample registration, chain-of-custody, full traceability
  • QC on every sample — RNA quality assessment (RNAscope), histopathological review by consulting pathologists
  • Scalable processing capacity — standardized sectioning, batch controls, and pre-analytical workflows that minimize variability across cohorts

Data Generation & Delivery

Fast, cost-efficient data generation, AI-ready delivery

  • Production capacity built for speed — redundant on-site platforms (2× Xenium, 2× Ventana, 2× PhenoImager) enable parallel runs and fast turnaround
  • Cost-optimization strategies — sample multiplexing, fit-for-purpose panel design, and tiered service options to match your budget
  • Open-format datasets — OME-TIFF, AnnData, H5 — delivered via secure high-speed transfer, ready to ingest
  • Optional data cleaning and normalization — kickstart your analysis with pre-processed datasets
Our data cleaning & preprocessing services
Discuss a dataset project

Our production platforms

Six on-site platforms for AI-grade dataset production

The Explicyte team at their Bordeaux laboratory

Paul Marteau, PharmD (study director), Imane Nafia, PhD (CSO), Loïc Cerf, MSc (COO), Alban Bessede, PhD (founder, CEO), Jean-Philippe Guégan, PhD (CTO)

contact our team

Discuss your AI dataset project

Whether you're training a foundation model, benchmarking a new approach, or building a validation cohort, we can scope the right dataset and propose next steps.Within 3 business days, a PhD-level scientist will reach out to discuss your project, propose the right cohort and dataset strategy, and provide pricing.

Answers about spatial biology datasets for AI

Frequently asked questions

What is AI-ready spatial biology data?

AI-ready spatial biology data refers to large-scale, structured datasets generated from human tissues — combining spatial transcriptomics, single-cell, or digital pathology readouts with clinical and pathological metadata — designed to be ingested directly into machine learning pipelines. Three properties matter most: consistent generation under quality-controlled conditions (to minimize batch effects), open-format delivery with documented schemas (to avoid preprocessing overhead), and verifiable provenance (to meet consent and regulatory requirements for AI training data). Explicyte generates AI-ready datasets across Xenium, Visium HD, Chromium, IHC, and multiplex IF.

Capacity depends on platform and panel design. With two on-site Xeniums, two Ventana stainers, and two PhenoImager scanners running in parallel, we can sustain production-scale throughput across cohorts of hundreds to thousands of samples. Typical project timelines range from a few weeks for focused panels on small cohorts to a few months for large multi-modality datasets. We confirm capacity and turnaround during project scoping based on your specific dataset specifications.

Datasets are delivered in open, AI-pipeline-compatible formats: OME-TIFF for imaging data (digital pathology, multiplex IF, Xenium), AnnData and H5 for single-cell and spatial transcriptomic data, and standard tabular formats (CSV, Parquet) for metadata and tabular results. Every dataset comes with full documentation — schema descriptions, data dictionaries, QC metrics, and methodology notes — to enable direct ingestion into your machine learning workflows. Custom formats can be accommodated on request.

Yes. Datasets generated by Explicyte are sourced under documented consent frameworks at our partner biobanks, with anonymized clinical and histopathological metadata — meeting the provenance and traceability requirements for AI training data. Specific data rights, exclusivity terms, and usage permissions are scoped per project. For sponsors with stricter data governance requirements (e.g., regulatory submissions, partner due diligence), we provide the documentation needed to demonstrate ethical sourcing throughout the data lifecycle.

All human tissues are sourced through partnerships with accredited French biobanks operating under NF S96-900 and ISO 20387 frameworks. Patient consent for research use is obtained at the biobank level; data reaches Explicyte fully anonymized. Chain-of-custody documentation is maintained throughout the project, and we can provide documentation packages for sponsor due diligence, regulatory review, or partner-required compliance frameworks (including the consent provenance increasingly required for AI training under emerging regulations like the EU AI Act).

Three things. First, we’re a wet lab partner, not a biospecimen vendor — sourcing, processing, and dataset generation happen end-to-end under one study director, so AI teams don’t need to coordinate across multiple suppliers. Second, we operate redundant on-site platforms (2× Xenium, 2× Ventana, 2× PhenoImager) for production-scale capacity and turnaround that single-instance labs can’t match. Third, we deliver in open formats (OME-TIFF, AnnData, H5) with documentation built for AI ingestion — proven through partnerships with foundation model developers and clinical research teams. ISO-certified digital pathology workflows and 10x Genomics Certified Service Provider status anchor the regulatory credibility.

Yes. Most of our AI partnerships involve project-specific cohort sourcing and dataset generation under exclusive-use terms. Data rights, exclusivity periods, and downstream usage permissions are negotiated as part of project scoping. We can also generate datasets where some metadata or derived outputs are co-owned or jointly publishable — flexible structures are possible depending on the project’s strategic goals.

Yes. Foundation model training requires datasets at production scale, with consistent generation conditions and clean, documented metadata — exactly what our redundant on-site platforms and ISO-certified workflows are built to deliver. We’ve supported foundation model partners through cohort design, sample sourcing under accredited biobank frameworks, multi-modality data generation (spatial transcriptomics + digital pathology), and dataset delivery in AI-pipeline-compatible formats.

Both. Sponsors who want raw outputs (FASTQ files, raw imaging data, unprocessed expression matrices) receive datasets ready for ingestion into their own pipelines. Sponsors who prefer pre-processed datasets can access optional bioinformatics preprocessing — quality control, normalization, cell segmentation, annotation — performed by our in-house data science team. The choice is scoped per project based on what fits your downstream workflow best.

Datasets are delivered via secure high-speed transfer in formats compatible with common cloud and on-premise infrastructure (AWS, GCP, Azure, on-prem clusters). For sponsors with specific pipeline requirements — direct cloud bucket delivery, custom schema mapping, or dataset versioning protocols — we can adapt delivery workflows during scoping. We do not require sponsors to use Explicyte-specific infrastructure or proprietary file formats at any stage.

Explicyte Oncology CRO logo

Capabilities

Modalities