Dataset Preview
Duplicate
The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
The dataset generation failed because of a cast error
Error code:   DatasetGenerationCastError
Exception:    DatasetGenerationCastError
Message:      An error occurred while generating the dataset

All the data files must have the same columns, but at some point there are 6 new columns ({'chr', 'start', 'gene', 'jaspar_id', 'tf', 'end'})

This happened while the csv dataset builder was generating data using

gzip://gene_cre_tf.tsv::hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz, ['hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/metadata/conditions.tsv', 'hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz']

Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1848, in _prepare_split_single
                  writer.write_table(table)
                  ~~~~~~~~~~~~~~~~~~^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 765, in write_table
                  self._write_table(pa_table, writer_batch_size=writer_batch_size)
                  ~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 773, in _write_table
                  pa_table = table_cast(pa_table, self._schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2378, in table_cast
                  return cast_table_to_schema(table, schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2306, in cast_table_to_schema
                  raise CastError(
                  ...<3 lines>...
                  )
              datasets.table.CastError: Couldn't cast
              gene: string
              chr: string
              start: int64
              end: int64
              tf: string
              jaspar_id: string
              condition: string
              -- schema metadata --
              pandas: '{"index_columns": [{"kind": "range", "name": null, "start": 0, "' + 1045
              to
              {'condition': Value('string')}
              because column names don't match
              
              During handling of the above exception, another exception occurred:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
                  parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
                                                                        ~~~~~~~~~~~~~~~~~~~~~~~~~^
                      builder, max_dataset_size_bytes=max_dataset_size_bytes
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
                  builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
                  ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1694, in _prepare_split
                  for job_id, done, content in self._prepare_split_single(
                                               ~~~~~~~~~~~~~~~~~~~~~~~~~~^
                      gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  ):
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1850, in _prepare_split_single
                  raise DatasetGenerationCastError.from_cast_error(
                  ...<4 lines>...
                  )
              datasets.exceptions.DatasetGenerationCastError: An error occurred while generating the dataset
              
              All the data files must have the same columns, but at some point there are 6 new columns ({'chr', 'start', 'gene', 'jaspar_id', 'tf', 'end'})
              
              This happened while the csv dataset builder was generating data using
              
              gzip://gene_cre_tf.tsv::hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz, ['hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/metadata/conditions.tsv', 'hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz']
              
              Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

condition
string
AD
APOE2vs3
LD-HighvsLow
APOE4vs2
APOE4vs3
APOEKOvs3
CD33
CLU-CRISPR
H1-IFN
HIV-ACTvsLAT
HIVactivated
HIVlatent
INPP5D
SORL1A528T
SORL1KO
TREM2KO
TREM2R47H
WTC11-IFN
iPSC-coculture
iPSC-xenot12d
iPSC-xenot7d
iPSC-xenot8w
SORL1KO
SORL1KO
SORL1KO
SORL1KO
SORL1KO
H1-IFN
SORL1KO
SORL1KO
H1-IFN
SORL1KO
SORL1KO
H1-IFN
SORL1KO
SORL1KO
SORL1KO
SORL1KO
SORL1KO
SORL1KO
CLU-CRISPR
CLU-CRISPR
CLU-CRISPR
iPSC-xenot8w
CD33
WTC11-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
CD33
WTC11-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
H1-IFN
WTC11-IFN
TREM2KO
SORL1A528T
SORL1KO
SORL1KO
SORL1KO
TREM2KO
WTC11-IFN
AD
H1-IFN
iPSC-xenot8w
SORL1KO
HIVactivated
AD
iPSC-xenot8w
SORL1KO
iPSC-xenot8w
SORL1KO
iPSC-xenot8w
SORL1KO
SORL1KO
H1-IFN
iPSC-xenot7d
SORL1KO
H1-IFN
H1-IFN
iPSC-xenot12d
iPSC-xenot7d
SORL1KO
iPSC-xenot8w
SORL1A528T
SORL1KO
H1-IFN
SORL1KO
H1-IFN
End of preview.

cEpiNets Dataset

Genome assembly: GRCh38 / hg38
Coordinate convention: 1-based, closed intervals
Cell type: human microglia and microglia-like cells

Processed Gene–CRE–TF network tables from cEpiNets (context-dependent epigenomic networks), the resource described in:

Fu, T.-T. et al. Context-dependent regulatory networks connect Alzheimer's disease genetics to microglial inflammatory states.

This repository contains the processed Gene–CRE–TF network data associated with cEpiNets. Interactive querying and visualization are available through the cEpiNets browser.

Overview

cEpiNets maps how transcription factors, cis-regulatory elements, and genes connect under different microglial contexts, including AD-related genetic perturbations and inflammatory or environmental conditions.

Each row in the core table represents one condition-specific Gene–CRE–TF association:

Gene — CRE — TF   (within one context)

The cEpiNets atlas described in the manuscript contains 191,292 CREs, 556 non-redundant JASPAR TF motif models, and more than 4 million TF–CRE links across 22 context pairs.

The downloadable table here is the filtered, gene-assigned network used by the public browser and therefore has a different row/CRE count from the full atlas:

gene_cre_tf.tsv.gz count
rows (Gene–CRE–TF–condition associations) 24,964,295
unique CREs 169,148
unique JASPAR motif models (jaspar_id) 556
unique TF names (tf) 529
unique genes / transcripts 28,837
conditions 22

The 556 jaspar_id values are the paper’s non-redundant JASPAR motif models. Multiple motif models can share the same TF name, so this table has 529 unique tf strings. CRE count is lower than 191,292 because this table retains the top 10% differential CREs after the 100 kb / 3-gene-per-CRE assignment used for the public network.

Data Files

path role size
network/gene_cre_tf.tsv.gz Core Gene–CRE–TF association table, 24,964,295 rows ~121 MB gzipped
metadata/conditions.tsv The 22 condition values in the network table
metadata/columns.md Column-level schema and coordinate conventions
LICENSE CC BY 4.0

Motif-level TF binding sites are not included. They do not contain gene names and are not required to use the network table.

File Schema

See metadata/columns.md. Summary of the core table:

column example
gene SORL1
chr chr20
start 30502691
end 30503583
tf EHF
jaspar_id MA0598.4
condition SORL1KO

The table is self-contained. Neo4j is not required to interpret a row.

Conditions

condition has 22 values, listed in metadata/conditions.tsv. Experimental design for each context is in the manuscript (Table S1).

Genome Assembly

All coordinates are GRCh38 / hg38, 1-based and closed (the start and end of a CRE are included). They are taken from the source chr:start-end interval strings without conversion to BED (0-based, half-open) coordinates.

Usage

import pandas as pd

net = pd.read_csv(
    "network/gene_cre_tf.tsv.gz",
    sep="\t",
    dtype={"start": "int32", "end": "int32"},
)

sorl1 = net[net["gene"] == "SORL1"]
trem2 = net[net["condition"] == "TREM2KO"]

From the Hugging Face Hub:

from huggingface_hub import hf_hub_download
import pandas as pd

path = hf_hub_download(
    repo_id="Thresh514/cEpiNets",
    filename="network/gene_cre_tf.tsv.gz",
    repo_type="dataset",
)
df = pd.read_csv(path, sep="\t")

Citation

Please cite the manuscript (DOI to be added on publication):

Fu, T.-T., Kurkela, M., Tu, J., Zhang, J., Sun, N., Farrer, L., TCW, J., and Hou, L.
Context-dependent regulatory networks connect Alzheimer's disease genetics
to microglial inflammatory states.

https://huggingface.co/datasets/Thresh514/cEpiNets

Correspondence: Lei Hou (leihou@bu.edu) and Julia TCW (juliatcw@bu.edu).

License

This dataset is released under Creative Commons Attribution 4.0 International (CC BY 4.0). See LICENSE.

Downloads last month
42