Skip to content

How to use a pre-trained scEmbed model

One advantage of scEmbed is the ability to use pre-trained models. This is useful for quickly getting embeddings of new data without having to train a new model. In this tutorial, we will show how to use a pre-trained model to get embeddings of new data.

I will be using the databio/luecken2021 model. It was trained on the Luecken2021 dataset, a first-of-its-kind multimodal benchmark dataset of 120,000 single cells from the human bone marrow of 10 diverse donors measured with two commercially-available multi-modal technologies: nuclear GEX with joint ATAC, and cellular GEX with joint ADT profiles.

This model will work best on PBMC-like data. It also requires your fragments be aligned to the GRCh38 genome.

Grab a fresh set of PBMC data from 10X genomics: https://www.10xgenomics.com/resources/datasets/10k-human-pbmcs-atac-v2-chromium-controller-2-standard

You need the Peak by cell matrix (filtered). This contains the binary accessibility matrix, the peaks, and the barcodes. Pre-trained models also requires that the data be in a scanpy.AnnData format and the .var attribute contain chr, start, and end values. For details on how to make this, see data preparation.

Once your data is ready, you can load it into python and get embeddings.

Encoding cells is as easy as:

import scanpy as sc
from geniml.scembed import ScEmbed
adata = sc.read_h5ad("path/to/adata.h5ad")
model = ScEmbed("databio/luecken2021")
embeddings = model.encode(adata)
adata.obsm['scembed_X'] = embeddings

And, thats it! You can now cluster your cells using the scembed_X embeddings.