#### ---- Description This dataset contains single-cell RNA-seq, surface protein, and TCR sequencing data generated from T cells using the 10x Genomics 5' kit. Both datasets were processed using Cell Ranger version 6.1.2 (10x Genomics). A pretrained variational autoencoder (VAE) model is also provided for detecting CD4 CTLs. #### ---- Revision history 2026-09-07 Dataset1.tar.gz unchanged Dataset2.tar.gz UPDATED - filtered_feature_bc_matrix/ (barcodes, features, matrix) replaced; the initial release contained Dataset1's files by mistake. VAE_model.tar.gz unchanged #### ---- Extract the archives tar -zxvf Dataset1.tar.gz tar -zxvf Dataset2.tar.gz tar -zxvf VAE_model.tar.gz #### ---- Dataset1.tar.gz ## T cells from 28 individuals (8 controls, 10 centenarians, 10 supercentenarians) Dataset1/ ├── cell.sample.cluster.43584.list (cell ID, sample ID, cluster name) ├── filtered_feature_bc_matrix/ │ ├── barcodes.tsv.gz │ ├── features.tsv.gz │ └── matrix.mtx.gz ├── vdj_t/ ├── clonotypes.csv └── filtered_contig_annotations.csv #### ---- Dataset2.tar.gz ## T cells from 6 individuals with and without PMA+ionomycin stimulation (2 controls, 2 centenarians, 2 supercentenarians) Dataset2/ ├── cell.sample.cluster.6400.list (cell ID, sample ID, cluster name) ├── filtered_feature_bc_matrix/ │ ├── barcodes.tsv.gz │ ├── features.tsv.gz │ └── matrix.mtx.gz ├── vdj_t/ ├── clonotypes.csv └── filtered_contig_annotations.csv #### ---- VAE_model.tar.gz ## Pre-trained VAE model and Python script for inference VAE_model/ ├── vae.py ├── 90.model/ │     ├── 01.gene.list │     └── 02.model.best.weights.h5 └── 91.input/ (Test Data)       ├── barcodes.tsv.gz       ├── features.tsv.gz       └── matrix.mtx.gz #### ---- Example: Loading Dataset1 into Seurat Seurat R package (version 4.1.1) in R (version 4.1.3) ## Load Seurat Package library(Seurat) ## Define paths dir <- "Dataset1/filtered_feature_bc_matrix/" cell <- "Dataset1/cell.sample.cluster.43584.list" ## Load RNA and protein data data <- Read10X(data.dir = dir) ## Remove Hashtag antibodies (used for sample multiplexing; not used for clustering) hashtags <- c("H1_TotalSeqC", "H3_TotalSeqC", "H5_TotalSeqC") data[["Antibody Capture"]] <- data[["Antibody Capture"]][!rownames(data[["Antibody Capture"]]) %in% hashtags, ] ## Create Seurat Object tc <- CreateSeuratObject(counts = data$`Gene Expression`) tc[['Protein']] <- CreateAssayObject(counts = data$`Antibody Capture`) ## Load metadata and filter cells used in downstream analysis tbl <- read.table("Dataset1/cell.sample.cluster.43584.list", row.names=1, col.names=c("bc","id","cluster")) tc <- subset(tc, cells = rownames(tbl)) tc <- AddMetaData(tc, metadata = tbl) #### ---- Example: Running VAE Inference ## Prerequisites: TensorFlow 2.20.0 or higher is required. ## (Using older versions will cause weight loading errors.) ## 1. Environment Setup conda create -n vae python=3.10 -y conda activate vae pip install --upgrade pip pip install "tensorflow[and-cuda]" pip install numpy pandas scikit-learn scipy ## 2. Version Check python -c "import tensorflow as tf; print(tf.__version__)" ## 3. Run Inference cd VAE_model python vae.py 91.input ## 4. Output Files ## Results are written to: ## 92.output/ ## ├── 21.valus.infer.txt (log file) ## ├── 22.laten.infer.tsv (16-dimensional latent space) ## └── 31.label.infer.tsv (Cell type prediction results: CD4_CTL, non_T, other_T)