import os
print(os.getcwd())
print(os.listdir())
import simicpipeline
print(f"SimiC pipeline version: ", {simicpipeline.__version__})/home/workdir
['data']
SimiC pipeline version: {'0.1.0'}
Author: Irene Marín-Goñi, PhD student - ML4BM group (CIMA University of Navarra)
This notebook demonstrates how to preprocess single-cell RNA-seq data for SimiC analysis.
This preprocessing tutorial covers: 1. Package installation and setup 2. MAGIC imputation pipeline 3. Gene selection and experiment setup 4. Preparing input files for SimiC
For running SimiC analysis see Tutorial_SimiCPipeline_full.ipynb
Before running SimiC, you need to: - Impute your scRNA-seq data. We recommend to use MAGIC and include a wrapper class MagicPipeline to ease the process. - Select top variable genes based on Median Absolute Deviation (MAD) or the genes of interest from which you want to infer the gene regulatory network. - Prepare input files in the correct format for SimiCPipeline.
This tutorial shows you how to do all of this using the SimiC preprocessing modules.
The easiest way to configure your environment is to follow the README instructions using poetry (or Docker).
Required packages for this tutorial: - simicpipeline - anndata - pandas - numpy - os - pickle
Internally simicpipeline also uses: - scipy - sklearn - scprep (in preprocessing) - magic-impute (in preprocessing)
First, import the necessary preprocessing modules.
/home/workdir
['data']
SimiC pipeline version: {'0.1.0'}
# Part 1: MAGIC Imputation Pipeline
MAGIC (Markov Affinity-based Graph Imputation of Cells) is used to denoise and impute scRNA-seq data. MagicPipeline facilitates the steps described in Magic Tutorial
Load your raw expression data. Note that MacigPipeline class expects AnnData format.
Below you will see different examples on how to generate the AnnData object from different input files (including Seurat if you are more familiar with R)
Note 1: If you have already filtered and transformed your data but want to repeat following these steps, make sure the adata object has the raw counts in the adata.raw.X slot
Note 2: If you have already processed (inputed) your data you can jump to Part 2 Experiment Setup
If you have your data in a Seurat object (R package) the easiest approach is this:
write.table(data.frame(Cells = colnames(seurat_obj)), file.path("/path/to/data/", "cell_ids.txt"), row.names = FALSE, col.names = FALSE, quote = FALSE)
write.table(data.frame(Genes = rownames(seurat_obj)), file.path("/path/to/data/", "genes_ids.txt"), row.names = FALSE, col.names = FALSE, quote = FALSE)
write.table(seurat_obj@metadata, file.path("/path/to/data/", "metadata.csv"), sep = ",",row.names = TRUE, col.names = TRUE, quote = FALSE)
m_raw = GetAssayData(seurat_obj, assay = "RNA", layer = "counts") # Take the raw counts
Matrix::writeMM(m_raw, paste0(magic_path, "/singlecell_matrix.mtx")) # Save it in MatrixMarket format Then follow the example code below.
print("Load your AnnData object here")
# # Example: Load from Matrix Market format
import pandas as pd
import anndata as ad
from pathlib import Path
# This will return a pd.DataFrame
df = simicpipeline.load_from_matrix_market(
matrix_path=Path("./data/simic_matrix.mtx"),
genes_path=Path("./data/simic_genes.txt"),
cells_path=Path("./data/simic_cells.txt"),
transpose=True,
cells_index_name="Cell",
)
adata = ad.AnnData(X=df.values, obs=pd.DataFrame(index=df.index), var=pd.DataFrame(index=df.columns))
obs_meta = pd.read_csv("./data/metadata.tsv", sep = "\t", index_col=0)
print(obs_meta.shape)
print(obs_meta.head())Load your AnnData object here
(72650, 8)
sample treatment cell_line \
cell
43_01_92__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_01_73__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_01_94__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_02_41__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_02_56__s1 KPB25L-UV_Combination Combination KPB25L-UV
final_annotation final_annotation_functional \
cell
43_01_92__s1 Cancer cells Proliferating cells
43_01_73__s1 Cancer cells Basal-like
43_01_94__s1 Cancer cells Proliferating cells
43_02_41__s1 Endothelial cells Endothelial cells
43_02_56__s1 Cancer cells Basal-like
nn_majority_label nn_majority_frac flag_misplaced
cell
43_01_92__s1 Proliferating cells 1.0 False
43_01_73__s1 Basal-like 1.0 False
43_01_94__s1 Proliferating cells 1.0 False
43_02_41__s1 Endothelial cells 1.0 False
43_02_56__s1 Basal-like 0.9 False
WARNING!!: Make sure that the index of obs_meta matches the cell names in df.index. Otherwise it will lead to misalignment of metadata and expression data.
Alternative loading options
# print("Load your AnnData object here")
# # Example: Load from CSV files
# import pandas as pd
# import anndata as ad
# # This example assumes that in csv rows are cells and columns are genes
# expression_data = pd.read_csv('path/to/expression_data.csv', index_col = 0) # Column 1 as cell IDs
# # Match the observation metadata
# metadata = pd.read_csv('path/to/metadata.csv', index_col=0)
# metadata = metadata.loc[expression_data.index]
# adata = ad.AnnData(X=expression_data.values, obs=metadata)
# # Example: Load from 10X format
# adata = ad.read_10x_mtx('path/to/10x/directory')
# # Example: Load from h5ad file using simicpipeline function
# adata = simicpipeline.load_from_anndata('path/to/your/data.h5ad')
# If your data is raw, you should set it properly with
# print(hasattr(adata, 'raw'))
# adata.raw = adata.copy()Create a MAGIC pipeline instance: - input_data: Your AnnData object. If you run the full pipline starting from raw counts they should be in adata.raw.X - project_dir: Project directory where magic_outputdir will be created and files will be saved - magic_output_file: Filename for the imputed data (default: ‘magic_data_allcells_sqrt.pickle’) - filtered: Set to True if data is already filtered (low quality cells and genes) (default: False)
# This command will initialize the MAGIC pipeline and generate the output directory if it does not exist
from simicpipeline import MagicPipeline
magic_pipeline = MagicPipeline(
input_data= adata,
project_dir='./SimiCExampleRun',
magic_output_file='magic_imputed.pickle',
filtered=False
)
print(magic_pipeline)Creating project directory: SimiCExampleRun
MagicPipeline(
data = AnnData object with (n_obs × n_vars) = 72650 × 36774,
filtered = False,
imputed = False,
magic_data = None,
project_dir = 'SimiCExampleRun'
)
min_cells_per_gene: Minimum number of cells expressing a gene (default: 10) - min_umis_per_cell: Minimum total UMI counts per cell (default: 500)
Note: If your data was already filtered you can skip this step and set the flitered argument flag to True in the previous step.
Filtering cells and genes...
Before filtering: 72650 cells x 36774 genes
Keeping 27837/36774 genes (75.70%)
Keeping 72650/72650 cells (100.00%)
All cells pass the filter!
After filtering: 72650 cells x 27837 genes
MagicPipeline(
data = AnnData object with (n_obs × n_vars) = 72650 × 27837,
filtered = True,
imputed = False,
magic_data = None,
project_dir = 'SimiCExampleRun'
)
Perform library size normalization with scprep followed by square root transformation.
Note: this will overide adata.X with normalized data and remove adata.raw slot.
Normalizing data...
After normalization: 72650 cells x 27837 genes
MagicPipeline(
data = AnnData object with (n_obs × n_vars) = 72650 × 27837,
filtered = True,
imputed = False,
magic_data = None,
project_dir = 'SimiCExampleRun'
)
Run magic imputation with defaul parameters
Running MAGIC imputation...
Calculating MAGIC...
Running MAGIC on 72650 cells and 27837 genes.
Calculating graph and diffusion operator...
Calculating PCA...
Calculated PCA in 39.74 seconds.
Calculating KNN search...
Calculated KNN search in 2.47 seconds.
Calculating affinities...
Calculated affinities in 5.33 seconds.
Calculated graph and diffusion operator in 47.65 seconds.
Running MAGIC with `solver='exact'` on 27837-dimensional data may take a long time. Consider denoising specific genes with `genes=<list-like>` or using `solver='approximate'`.
Calculating imputation...
Calculated imputation in 169.16 seconds.
Calculated MAGIC in 218.13 seconds.
MAGIC imputation complete: 72650 cells x 27837 genes
Saving MAGIC-imputed data to SimiCExampleRun/magic_output/magic_imputed.pickle
Saved successfully to SimiCExampleRun/magic_output/magic_imputed.pickle
MagicPipeline(
data = AnnData object with (n_obs × n_vars) = 72650 × 27837,
filtered = True,
imputed = True,
magic_data = AnnData object with n_obs × n_vars = 72650 × 27837,
project_dir = 'SimiCExampleRun'
)
If you want to run MAGIC imputation with custom parameters you can pass them as **kwargs: - t: Number of diffusion steps (default: ‘auto’) - knn: Number of nearest neighbors (default: 5) - decay: Decay rate for kernel (default: 1) - n_jobs: Number of parallel jobs (default: -2) - genes: Genes to be returned. If None or “all genes” it returns teh entire matrix. - save_data: Whether to automatically save imputed data (default: True). If magic_output_file extension is .pickle will save it in .pickle, if h5ad, will save in adata format.
See MAGIC documentation for more parameter options.
SimiCExampleRun/
└── magic_output/
├── magic_imputed.h5ad
└── magic_imputed.pickle
Success! MAGIC imputation is complete. The imputed data is saved in the magic_output directory.
Note: if you have your data filtered, normalized and imputed with alternative methods you can start from here.
In this example we will start from the imputed AnnData object from the MAGIC pipeline. If you saved and stopped your work, you can re-load the object with the following code:
AnnData object with n_obs × n_vars = 72650 × 27837
obs: 'sample', 'treatment', 'cell_line', 'final_annotation', 'final_annotation_functional', 'nn_majority_label', 'nn_majority_frac', 'flag_misplaced'
Imputed data shape: (72650, 27837)
sample treatment cell_line \
Cell
43_01_73__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_01_92__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_01_94__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_02_41__s1 KPB25L-UV_Combination Combination KPB25L-UV
43_02_56__s1 KPB25L-UV_Combination Combination KPB25L-UV
final_annotation final_annotation_functional \
Cell
43_01_73__s1 Cancer cells Basal-like
43_01_92__s1 Cancer cells Proliferating cells
43_01_94__s1 Cancer cells Proliferating cells
43_02_41__s1 Endothelial cells Endothelial cells
43_02_56__s1 Cancer cells Basal-like
nn_majority_label nn_majority_frac flag_misplaced
Cell
43_01_73__s1 Basal-like 1.0 False
43_01_92__s1 Proliferating cells 1.0 False
43_01_94__s1 Proliferating cells 1.0 False
43_02_41__s1 Endothelial cells 1.0 False
43_02_56__s1 Basal-like 0.9 False
Create an experiment setup instance and directories: - input_data: Your imputed AnnData object or pandas DataFrame (cells × genes) - tf_path: Path to transcription factor (TF) list file (.csv or .txt) - project_dir: Directory where experiment files will be saved
Note: In case you do not have a TF list:
We provide a mouse TF list in the data folder that can be saved in your working data directory. - TF mouse list was downloaded in December 2024 from AnimalTFDB4 - TF human list was downloaded in February 2026 from XX
In this tutorial we are working wiht mouse data so we will use the TF list from AnimalTFDB4
# Initialize ExperimentSetup
from simicpipeline import ExperimentSetup
experiment = ExperimentSetup(
input_data = imputed_data,
tf_path = "./data/TF_list.csv", # Should have no header
project_dir='./SimiCExampleRun'
)
print(f"Matrix shape: {experiment.matrix.shape}")
print(f"Number of cells: {len(experiment.cell_names)}")
print(f"Number of genes: {len(experiment.gene_names)}")
print(f"Number of TFs: {len(experiment.tf_list)}")
print(f"... Example TF names: {experiment.tf_list[0:5]}\n")
print("\n" + "="*70)
print(f"Current directory status")
print("="*70 + "\n")
experiment.print_project_info(max_depth=2)Matrix shape: (72650, 27837)
Number of cells: 72650
Number of genes: 27837
Number of TFs: 1611
... Example TF names: ['Lin28b', 'Tbx2', 'Dmtf1', 'Irx4', 'Irf3']
======================================================================
Current directory status
======================================================================
SimiCExampleRun/
├── inputFiles/
├── magic_output/
│ ├── magic_imputed.h5ad
│ └── magic_imputed.pickle
└── outputSimic/
├── figures/
└── matrices/
The previous code automatically creates the SimiC directory structure:
project_dir/
├── inputFiles/ # Input files for SimiC
└── outputSimic/ # Output files generated by SimiCPipeline
├── figures/ # For future visualizations
└── matrices/ # For future results
Select top variable genes based on Median Absolute Deviation (MAD): - n_tfs: Number of top TF genes to select (default: 100) - n_targets: Number of top target genes to select (default: 1000)
Returns a tuple of (TF_list, TARGET_list)
Removing 0 targets with MAD = 0
Selecting top 1000 targets based on MAD.
Selected 100 TFs
Selected 1000 targets
Top 10 TFs: ['Hmga2', 'Zfpm2', 'Bnc2', 'Glis3', 'Zeb2', 'Pbx1', 'Ebf1', 'Zeb1', 'Mecom', 'Nfib']
Top 10 targets: ['Rn18s-rs5', 'Malat1', 'Cmss1', 'Lamp2', 'Brinp3', 'Rad51b', 'Ccbe1', 'Tenm4', 'Nop58', 'Xist']
Create a subset of your data containing only the selected TFs and targets.
# Combine TF and target lists
import anndata as ad
selected_genes = tf_list + target_list
# Subset the data
if isinstance(imputed_data, ad.AnnData):
subset_data = imputed_data[:, selected_genes].copy()
elif isinstance(imputed_data, pd.DataFrame):
subset_data = imputed_data[selected_genes].copy()
print(f"Subset data shape: {subset_data.shape}")Subset data shape: (72650, 1100)
Save the expression matrix and TF names in .pickle format and annotation file (optional) as .txt - run_data: ad.AnnData or pd.Dataframe with data to run in SimiC (Inputed and sliced according to experiment run) - matrix_filename: Filename to save run_data (saved with row/column headers). Can be .pickleor csv. - tf_filename: Filename for TF names list for the experiment run. Can be .pickleor csv. Even though you have a general TF_list file, this function will save the TFs selected by MAD that are found in your run_data.
annotation:str (Optional) if run_data is ad.AnnData and annotation is in run_data.obs.columns, it will create a .csv file with the phenotype annotations needed for SimiC with cell names as index and columns category and labels.
annotation_order(Optional) List defining the desired order of annotation categories (e.g. [‘control’, ‘treated’]). The first element maps to 0, second to 1, etc. If None, pd.factorize default order is used.
All files are saved in the inputFiles/ directory.
Saved expression matrix to SimiCExampleRun/inputFiles/expression_matrix.csv
Saved 100 TFs to SimiCExampleRun/inputFiles/TF_list.csv
Warning: Annotation 'groups' not found in obs columns.
Available columns:
['sample', 'treatment', 'cell_line', 'final_annotation', 'final_annotation_functional', 'nn_majority_label', 'nn_majority_frac', 'flag_misplaced']
Please manually provide an appropriate annotation file to SimiCPipeline in SimiCExampleRun/inputFiles
-------
Experiment files saved successfully.
-------
We recommend saving it in pickle format for fast load/dump process and save disk space.
Regarding label handling. Because it is important for SimiCPipeline run to clearly define the desired labels in the correct order, several checks and warnings will be raised if annotation_order is wrongly provided.
experiment.save_experiment_files(
run_data = subset_data,
matrix_filename = 'expression_matrix.pickle',
tf_filename = 'TF_list.csv', # Will raise warning if file already exists and overwrite
annotation = 'treatment', # Will rase WARNING because it is categorical, it will convert it to numerical and assign the order with the annotation_order argument
annotation_order = ['control', 'DAC'] # Will raise ERROR if missing categories or if categories in annotation_order do not match those in subset_data.obs['treatment']
)Saved expression matrix to SimiCExampleRun/inputFiles/expression_matrix.pickle
Warning: Output file SimiCExampleRun/inputFiles/TF_list.csv already exists and will be overwritten.
Saved 100 TFs to SimiCExampleRun/inputFiles/TF_list.csv
-------
Annotation 'treatment' found in obs columns!
Warning: annotation is not numeric. Will convert from categorical to numeric.
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[21], line 1 ----> 1 experiment.save_experiment_files( 2 run_data = subset_data, 3 matrix_filename = 'expression_matrix.pickle', 4 tf_filename = 'TF_list.csv', # Will raise warning if file already exists and overwrite 5 annotation = 'treatment', # Will rase WARNING because it is categorical, it will convert it to numerical and assign the order with the annotation_order argument 6 annotation_order = ['control', 'DAC'] # Will raise ERROR if missing categories or if categories in annotation_order do not match those in subset_data.obs['treatment'] 7 ) File /home/SimiCPipeline/src/simicpipeline/core/simicpreprocess.py:540, in ExperimentSetup.save_experiment_files(self, run_data, matrix_filename, tf_filename, annotation, annotation_order) 538 unused_in_order = ordered_set - observed_categories 539 if missing_from_order: --> 540 raise ValueError(f"The following annotation values are in the data but not in annotation_order: {missing_from_order}.\n" 541 "If only these annotation_order categories are needed, subset run_data accordingly before running `save_experiment_files`" 542 ) 543 if unused_in_order: 544 raise ValueError( 545 f"The following annotation_order values are not present in data: {unused_in_order}" 546 ) ValueError: The following annotation values are in the data but not in annotation_order: {'Combination', 'PD-L1'}. If only these annotation_order categories are needed, subset run_data accordingly before running `save_experiment_files`
Warning: Output file SimiCExampleRun/inputFiles/expression_matrix.pickle already exists and will be overwritten.
Saved expression matrix to SimiCExampleRun/inputFiles/expression_matrix.pickle
Warning: Output file SimiCExampleRun/inputFiles/TF_list.csv already exists and will be overwritten.
Saved 100 TFs to SimiCExampleRun/inputFiles/TF_list.csv
-------
Annotation 'treatment' found in obs columns!
Warning: annotation is not numeric. Will convert from categorical to numeric.
Annotation order applied: {0: 'control', 1: 'PD-L1', 2: 'DAC', 3: 'Combination'}
Annotation distribution:
label
3 17259
2 20491
1 15426
0 19474
Saved annotation to SimiCExampleRun/inputFiles/treatment_annotation.csv
-------
Experiment files saved successfully.
-------
Be careful! Double check your labels are correct and match your cell numbers
Success! All preprocessing steps completed. Your files are ready for SimiC analysis.
Check that all files were created correctly.
This tutorial covered:
✓ Loading and filtering scRNA-seq data
✓ Running MAGIC imputation
✓ Selecting top variable genes using MAD
✓ Preparing input files for SimiC analysis with proper directory structure
Your output directory now contains:
SimicExampleRun/
├── magic_output/
│ └── magic_imputed.pickle/.h5ad
├── inputFiles/
│ ├── expression_matrix.pickle/.csv
│ ├── TF_list.pickle
│ └── treatment_annotation.csv
└── outputSimic/
├── figures/
└── matrices/
In this the previous section we used the whole Magic-inputed matrix (obtained in Part1) and selected top MAD genes but you may want to run SimiC in a subset of cells from your data.
Generally we recommend to impute the data in the whole dataset, especially if it was generated in the same sequencing batch, as MAGIC will have more context information to impute the data. However, we acknowledge that every experiment/dataset is different and may require a different approach.
If you want to run SimiCPipeline in a subset of cells, once you have imputed your data, make sure you slice the adata object before you inilitalize the ExperimentSetup class so MAD genes are calculated over your cells of interest.
Following this tutorial steps are not required for running SimiCPipeline but recommended before as it will facilitate the process. Just make sure that:
We will show how to easily subset your cells and save it with the ExperimentSetup
| sample | treatment | cell_line | final_annotation | final_annotation_functional | nn_majority_label | nn_majority_frac | flag_misplaced | |
|---|---|---|---|---|---|---|---|---|
| Cell | ||||||||
| 43_01_73__s1 | KPB25L-UV_Combination | Combination | KPB25L-UV | Cancer cells | Basal-like | Basal-like | 1.0 | False |
| 43_01_92__s1 | KPB25L-UV_Combination | Combination | KPB25L-UV | Cancer cells | Proliferating cells | Proliferating cells | 1.0 | False |
| 43_01_94__s1 | KPB25L-UV_Combination | Combination | KPB25L-UV | Cancer cells | Proliferating cells | Proliferating cells | 1.0 | False |
| 43_02_41__s1 | KPB25L-UV_Combination | Combination | KPB25L-UV | Endothelial cells | Endothelial cells | Endothelial cells | 1.0 | False |
| 43_02_56__s1 | KPB25L-UV_Combination | Combination | KPB25L-UV | Cancer cells | Basal-like | Basal-like | 0.9 | False |
| ... | ... | ... | ... | ... | ... | ... | ... | ... |
| 06_92_89__s8 | KPB25L_control | control | KPB25L | Cancer cells | Basal-like | Basal-like | 1.0 | False |
| 06_94_61__s8 | KPB25L_control | control | KPB25L | Macrophages | Macrophages | Macrophages | 1.0 | False |
| 06_96_48__s8 | KPB25L_control | control | KPB25L | Cancer cells | Unknown | Unknown | 1.0 | False |
| 06_96_58__s8 | KPB25L_control | control | KPB25L | Macrophages | Macrophages | Macrophages | 1.0 | False |
| 06_96_82__s8 | KPB25L_control | control | KPB25L | Cancer cells | Basal-like | Basal-like | 1.0 | False |
72650 rows × 8 columns
from simicpipeline import ExperimentSetup
cell_mask = imputed_data.obs['cell_line'].isin(["KPB25L"]) & imputed_data.obs['final_annotation_functional'].isin(['Proliferating cells','Basal-like'])
print(f"Number of selected cells:",{sum(cell_mask)})
subset_imputed_data = imputed_data[cell_mask,:].copy()
print(f"Suset matrix shape:", {subset_imputed_data.shape})
experiment2 = ExperimentSetup(
input_data = subset_imputed_data,
tf_path = "./data/TF_list.csv", # Should have no header
project_dir='./SimiCExampleRun/KPB25L/Tumor'
)
experiment2.print_project_info(max_depth=1)
# Then follow the same steps as above to select genes and save experiment filesNumber of selected cells: {21490}
Suset matrix shape: {(21490, 27837)}
Creating project directory: SimiCExampleRun/KPB25L/Tumor
Tumor/
├── inputFiles/
└── outputSimic/
However the initial output directory will then look like:
SimiCExampleRun/
├── KPB25L/
│ └── Tumor/
│ ├── inputFiles/
│ └── outputSimic/
├── inputFiles/
│ ├── TF_list.csv
│ ├── expression_matrix.csv
│ ├── expression_matrix.pickle
│ └── treatment_annotation.csv
├── magic_output/
│ ├── magic_imputed.h5ad
│ └── magic_imputed.pickle
└── outputSimic/
├── figures/
└── matrices/
Repeat the steps to calculate MAD genes adn save files
tf_list, target_list = experiment2.calculate_mad_genes(
n_tfs=100,
n_targets=1000
)
# Combine TF and target lists
selected_genes2 = tf_list + target_list
subset_experiment_data = subset_imputed_data[:, selected_genes2].copy()
experiment2.save_experiment_files(
run_data = subset_experiment_data,
matrix_filename = 'expression_matrix.pickle',
tf_filename = 'TF_list.csv', # Will raise warning if file already exists and overwrite
annotation = 'treatment',
annotation_order = ['control', 'PD-L1','DAC', 'Combination']
)
print(f"Subset data shape: {subset_experiment_data.shape}")Removing 8 targets with MAD = 0
Selecting top 1000 targets based on MAD.
Saved expression matrix to SimiCExampleRun/KPB25L/Tumor/inputFiles/expression_matrix.pickle
Saved 100 TFs to SimiCExampleRun/KPB25L/Tumor/inputFiles/TF_list.csv
-------
Annotation 'treatment' found in obs columns!
Warning: annotation is not numeric. Will convert from categorical to numeric.
Annotation order applied: {0: 'control', 1: 'PD-L1', 2: 'DAC', 3: 'Combination'}
Annotation distribution:
label
3 4392
2 5895
1 4964
0 6239
Saved annotation to SimiCExampleRun/KPB25L/Tumor/inputFiles/treatment_annotation.csv
-------
Experiment files saved successfully.
-------
Subset data shape: (21490, 1100)
SimiCPipeline class to run SimiC.SimicVisualization class to analyze GRNs and TF activities.Check Tutorial_SimiCPipeline_full.ipynb or Tutorial_SimiCPipeline_visualization for guided info.
Data Format: All matrices are stored as cells × genes (rows = cells, columns = genes)
Memory Usage: MAGIC imputation can be memory-intensive for large datasets. Consider using a machine with sufficient RAM and adjusting MAGIC parameters (n_jobs, knn, t)
Please note: Although you will be able to pass custom file/direcotry paths, we highly recommend to follow the directory structure described above and follow this tutorial before running SimiC to avoid errors.