Pando Bioscience, AI enzyme design
Database

Pando Enzyme Database

The knowledge behind our AI. 22 billion high-quality sequences — 150× larger than UniProtKB — purpose-encoded for AI enzyme discovery.

By the numbers
0B

high‑quality, well‑annotated sequences

0.0B

non‑redundant sequences at 50% identity cutoff

0×

larger than UniProtKB

Database growth

Pando grows as its database grows

Proprietary sequence data, growing year over year.

Bar chart of database size: 0.2 billion sequences in 2023 (UniProtKB baseline), 0.6 in 2024, 1.2 in 2025, 3.9 in February 2026, 22 billion in September 20260B1B2B3B4B11B22B0.2BUniProtKB20230.6B20241.2B20253.9BFeb 202622BSep 2026
How are we using it?

A database built for discovery and model training

The Pando database was built to find — every sequence is cleaned, quality-controlled, annotated, and encoded so AI can search it and learn from it.

1

Patent-derived sequences included

Sequences mined from the patent literature, capturing decades of industrial enzyme engineering that public archives miss.

2

Metagenomes from extreme habitats

Fresh metagenomic diversity from extreme environments — the natural reservoir of stability and new chemistry.

3

Rich metadata and functional annotations

Every entry carries rich metadata and functional annotations, so models learn function — not just letters.

4

Encoded for AI enzyme discovery

Pre-encoded representations purpose-built for AI search, ranking, and generative design.

Enzyme universe

Surfing the enzyme universe

A schematic map of protein space, billions of proteins strong and encoded for AI search.

Enzyme libraries

An extensive collection of enzyme libraries

We have an extensive collection of enzyme libraries, including:

ssRNA ligasesdsDNA ligaseskinasesKREDsIREDshydrataseshydrolasesoxidasesamidasescarboxylasesdecarboxylasestransferasesglutathione synthetaseDNA polymerasesreverse transcriptasesUGTsand many more
Get in touch

Want to know more?

Let’s explore the protein universe together.