Pando Enzyme Database
The knowledge behind our AI. 22 billion high-quality sequences — 150× larger than UniProtKB — purpose-encoded for AI enzyme discovery.
high‑quality, well‑annotated sequences
non‑redundant sequences at 50% identity cutoff
larger than UniProtKB
Pando grows as its database grows
Proprietary sequence data, growing year over year.
A database built for discovery and model training
The Pando database was built to find — every sequence is cleaned, quality-controlled, annotated, and encoded so AI can search it and learn from it.
Patent-derived sequences included
Sequences mined from the patent literature, capturing decades of industrial enzyme engineering that public archives miss.
Metagenomes from extreme habitats
Fresh metagenomic diversity from extreme environments — the natural reservoir of stability and new chemistry.
Rich metadata and functional annotations
Every entry carries rich metadata and functional annotations, so models learn function — not just letters.
Encoded for AI enzyme discovery
Pre-encoded representations purpose-built for AI search, ranking, and generative design.
Surfing the enzyme universe
A schematic map of protein space, billions of proteins strong and encoded for AI search.
An extensive collection of enzyme libraries
We have an extensive collection of enzyme libraries, including:
Want to know more?
Let’s explore the protein universe together.
