The Catalog

Broad chemistry, optimized for drug-like characteristics

What the collection actually looks like as chemical matter, measured end to end.

Most screening libraries are assembled. Ours is generated, filtered, and scored, which means we can say precisely what is in it and show the measurements behind every claim. This page is that description: what the catalog covers, how it is distributed, and where it is deliberately different from a curated bioactive library.

More chemistry per molecule, in a tighter window

The catalog occupies a deliberately narrow region of chemical space, and within that region it carries substantially more structural variety per molecule than a curated bioactive library.

Every comparison below was measured against 771,668 ChEMBL compounds using the same fingerprint and the same measurement code for both collections. Where the two columns were computed over different-sized populations, the row says so, because sample size affects some of these quantities.

Structural diversity
MetricCatalogChEMBL reference
Unique scaffolds per moleculehigher means more structural variety per moleculeMeasured on a random 500,000-molecule sample0.420.29
Scaffolds appearing exactly onceMeasured on a random 500,000-molecule sample79.6%60.0%
Mean pairwise similaritylower means the collection spans a wider region of chemical spaceMeasured on every molecule in each collection0.1620.119
Median nearest-neighbour similarityhow close a typical molecule's closest relative isMeasured on 2,000 molecules each, searched against the entire collection: 11.2 million and 771,6680.7860.784
Morgan fingerprints, radius 2, 2048 bits, Tanimoto similarity.

Those last two rows say different things, and both matter. The higher mean pairwise similarity means the catalog sits inside a tighter envelope: pick two molecules at random and they are more alike than two random ChEMBL compounds, because the property filters keep everything in lead-like territory. The near-identical nearest-neighbour figures mean that despite that tighter envelope, a typical catalog molecule has a closest relative about as near as a ChEMBL molecule does.

Scale is what makes those compatible. The catalog is roughly fourteen times the size of the reference set, so even though its molecules are spread more thinly through chemical space, there are enough of them that any given molecule still has a close relative nearby. Measured at matched size the difference shows: on a 2.25 million subset, still far larger than the reference, the median nearest-neighbour similarity falls to 0.737.

For a primary screening campaign the combination is a good one. You get 1.4 times the scaffold variety per molecule, inside a property window already filtered for drug-likeness, without the collection being so sparse that individual molecules sit isolated.

How close is the nearest relative?

This is the share of molecules having at least one neighbour at each similarity threshold. Both columns are measured the same way: 2,000 randomly chosen molecules, each compared exhaustively against every other molecule in its collection. We use direct comparison rather than reading neighbours off the clustering, because clustering systematically overstates how isolated a molecule is.

Molecules with a neighbour at or above each threshold
MetricCatalogChEMBL reference
Tanimoto 0.4599.3%98.7%
Tanimoto 0.5595.6%97.1%
Tanimoto 0.6585.2%89.8%
Tanimoto 0.7562.3%65.8%
Tanimoto 0.8523.8%21.3%
Catalog: 11,165,044 molecules searched. Reference: 771,668.

The two collections track each other closely. Essentially every catalog molecule has a related neighbour, six in ten have a close one at 0.75, and at the tightest threshold the catalog carries slightly more near-identical pairs than the reference. Whatever structural variety the catalog gains, it does not come at the cost of leaving molecules stranded on their own.

A deliberately lead-like envelope

The catalog spans a narrower range than the reference on every property axis we measured, most visibly in molecular weight and ring count. Medians sit close to ChEMBL’s; the tails are cut.

That is a filtering decision, not an accident. Candidates outside drug-like property ranges are discarded, as are molecules that look impractical to synthesize and molecules flagged by medicinal-chemistry structural alerts for reactive groups or implausible ring systems. The result is a library where the property screen has already been applied.

Novel, and still growing into new chemistry

Our generative models are trained to be novel. Every candidate is checked against ChEMBL and other sources and discarded if found elsewhere.

Novelty and growth
MetricMeasured
Catalog molecules also found in ChEMBLexcluded by construction, and re-checked on every importMeasured on every molecule, on every import0%
Candidates discarded for matching a known ChEMBL compoundmeasured before filtering, so these never enter the catalogMeasured on one generation run of 699,785 candidates0.50%
Scaffolds in the newest generation run absent from all earlier runsMeasured on 100,000 molecules from each of five generation eras64.7%
Additional molecules contributing a previously unseen scaffoldMeasured on a random 500,000-molecule sample~1 in 3
Largest single scaffold family, as a share of the catalogMeasured on a random 500,000-molecule sample1.85%

That second row is a statement about the generator, not about the catalog. The rate at which they re-emit a compound that already exists in the training data measures how much they memorized. At 0.50% it is roughly four times lower than the size of the two collections alone would predict, and every one of those matches is discarded before import.

The models learned the style of bioactive chemistry rather than memorizing its contents, and the collection is still reaching structures it has not covered before rather than repeating itself.

What this catalog is not

The catalog averages roughly 2.4 molecules per scaffold, and very close analogs are about half as common as in the reference library. It is not built to hand you a ready-made structure-activity series around a hit. A diverse library is optimized for finding starting points; a focused library is optimized for developing them, and those are different collections.

In practice that means follow-up around a confirmed hit is best served by us generating analogs on demand around that specific scaffold.

How we measured this

Every molecule in the catalog was processed, not a sample of it. Each was converted to a Morgan fingerprint (radius 2, 2048 bits, bit vector) using the same code path that powers similarity search in the product, so the numbers on this page describe the same representation the application uses.

Clustering

Molecules were grouped with BitBIRCH, a tree-based clustering method built for large molecular libraries: it scales linearly with library size and takes a single interpretable parameter, a Tanimoto similarity threshold, rather than a target number of clusters. We used the diameter criterion, so a molecule joins a group only if the group’s average pairwise similarity stays above the threshold. Groups were built at two thresholds, 0.45 and 0.65, because coverage and tightness trade off against each other and neither setting alone describes the collection.

A refinement pass follows the initial build. Tree-based clustering routes each molecule by comparing against group representatives, and in high-dimensional fingerprint space that misroutes: molecules get stranded alone even when a close relative is present.

The reference arm

Every ChEMBL comparison on this page was computed locally over 771,668 compounds, with the identical fingerprint and the identical code that processed our own molecules. We deliberately avoided quoting published figures for reference libraries: different fingerprints and different implementations produce different numbers, and a comparison assembled from two sources is not a comparison.

What is measured on what

Some figures come from the complete catalog and some from samples, and the distinction matters when reading them.

Measurement basis
MetricBasis
Molecule countLive query
Property distributions and Rule of FiveEvery molecule
ClusteringEvery molecule
Mean pairwise similaritycomputed from bit sums, so no sampling neededEvery molecule
Scaffold coveragestratified across the whole catalog500,000 sample
Nearest-neighbour distributioneach query exhaustively compared against all 2.25M molecules in one block2,000 queries
Property envelopedescriptors recomputed from structure250,000 sample
ChEMBL reference771,668 compounds

Samples are drawn across the full id range rather than from the beginning of the catalog. Nearest-neighbour figures come from exhaustive comparison against every other molecule, not from cluster membership, because clustering systematically over-reports isolation even after refinement.

Known limits

Scaffold counts use Bemis-Murcko perception, which treats a single ring-heteroatom change as a different scaffold. That makes our analog-depth numbers conservative, and it applies equally to both collections. Morgan fingerprints ignore stereochemistry, so stereoisomer pairs register as identical. All similarity figures use one fingerprint and one metric; other representations would shift the absolute values, though the catalog-versus-reference comparisons are computed the same way on both sides.