The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
Error code: DatasetGenerationCastError
Exception: DatasetGenerationCastError
Message: An error occurred while generating the dataset
All the data files must have the same columns, but at some point there are 6 new columns ({'chr', 'start', 'gene', 'jaspar_id', 'tf', 'end'})
This happened while the csv dataset builder was generating data using
gzip://gene_cre_tf.tsv::hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz, ['hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/metadata/conditions.tsv', 'hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz']
Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)
Traceback: Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1848, in _prepare_split_single
writer.write_table(table)
~~~~~~~~~~~~~~~~~~^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 765, in write_table
self._write_table(pa_table, writer_batch_size=writer_batch_size)
~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 773, in _write_table
pa_table = table_cast(pa_table, self._schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2378, in table_cast
return cast_table_to_schema(table, schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2306, in cast_table_to_schema
raise CastError(
...<3 lines>...
)
datasets.table.CastError: Couldn't cast
gene: string
chr: string
start: int64
end: int64
tf: string
jaspar_id: string
condition: string
-- schema metadata --
pandas: '{"index_columns": [{"kind": "range", "name": null, "start": 0, "' + 1045
to
{'condition': Value('string')}
because column names don't match
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
~~~~~~~~~~~~~~~~~~~~~~~~~^
builder, max_dataset_size_bytes=max_dataset_size_bytes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1694, in _prepare_split
for job_id, done, content in self._prepare_split_single(
~~~~~~~~~~~~~~~~~~~~~~~~~~^
gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
):
^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1850, in _prepare_split_single
raise DatasetGenerationCastError.from_cast_error(
...<4 lines>...
)
datasets.exceptions.DatasetGenerationCastError: An error occurred while generating the dataset
All the data files must have the same columns, but at some point there are 6 new columns ({'chr', 'start', 'gene', 'jaspar_id', 'tf', 'end'})
This happened while the csv dataset builder was generating data using
gzip://gene_cre_tf.tsv::hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz, ['hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/metadata/conditions.tsv', 'hf://datasets/Thresh514/cEpiNets@d6993b49601a4d40aba3bff0e3afe9277d9c0373/network/gene_cre_tf.tsv.gz']
Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
condition string |
|---|
AD |
APOE2vs3 |
LD-HighvsLow |
APOE4vs2 |
APOE4vs3 |
APOEKOvs3 |
CD33 |
CLU-CRISPR |
H1-IFN |
HIV-ACTvsLAT |
HIVactivated |
HIVlatent |
INPP5D |
SORL1A528T |
SORL1KO |
TREM2KO |
TREM2R47H |
WTC11-IFN |
iPSC-coculture |
iPSC-xenot12d |
iPSC-xenot7d |
iPSC-xenot8w |
SORL1KO |
SORL1KO |
SORL1KO |
SORL1KO |
SORL1KO |
H1-IFN |
SORL1KO |
SORL1KO |
H1-IFN |
SORL1KO |
SORL1KO |
H1-IFN |
SORL1KO |
SORL1KO |
SORL1KO |
SORL1KO |
SORL1KO |
SORL1KO |
CLU-CRISPR |
CLU-CRISPR |
CLU-CRISPR |
iPSC-xenot8w |
CD33 |
WTC11-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
CD33 |
WTC11-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
H1-IFN |
WTC11-IFN |
TREM2KO |
SORL1A528T |
SORL1KO |
SORL1KO |
SORL1KO |
TREM2KO |
WTC11-IFN |
AD |
H1-IFN |
iPSC-xenot8w |
SORL1KO |
HIVactivated |
AD |
iPSC-xenot8w |
SORL1KO |
iPSC-xenot8w |
SORL1KO |
iPSC-xenot8w |
SORL1KO |
SORL1KO |
H1-IFN |
iPSC-xenot7d |
SORL1KO |
H1-IFN |
H1-IFN |
iPSC-xenot12d |
iPSC-xenot7d |
SORL1KO |
iPSC-xenot8w |
SORL1A528T |
SORL1KO |
H1-IFN |
SORL1KO |
H1-IFN |
cEpiNets Dataset
Genome assembly: GRCh38 / hg38
Coordinate convention: 1-based, closed intervals
Cell type: human microglia and microglia-like cells
Processed Gene–CRE–TF network tables from cEpiNets (context-dependent epigenomic networks), the resource described in:
Fu, T.-T. et al. Context-dependent regulatory networks connect Alzheimer's disease genetics to microglial inflammatory states.
This repository contains the processed Gene–CRE–TF network data associated with cEpiNets. Interactive querying and visualization are available through the cEpiNets browser.
Overview
cEpiNets maps how transcription factors, cis-regulatory elements, and genes connect under different microglial contexts, including AD-related genetic perturbations and inflammatory or environmental conditions.
Each row in the core table represents one condition-specific Gene–CRE–TF association:
Gene — CRE — TF (within one context)
The cEpiNets atlas described in the manuscript contains 191,292 CREs, 556 non-redundant JASPAR TF motif models, and more than 4 million TF–CRE links across 22 context pairs.
The downloadable table here is the filtered, gene-assigned network used by the public browser and therefore has a different row/CRE count from the full atlas:
gene_cre_tf.tsv.gz |
count |
|---|---|
| rows (Gene–CRE–TF–condition associations) | 24,964,295 |
| unique CREs | 169,148 |
unique JASPAR motif models (jaspar_id) |
556 |
unique TF names (tf) |
529 |
| unique genes / transcripts | 28,837 |
| conditions | 22 |
The 556 jaspar_id values are the paper’s non-redundant JASPAR motif models. Multiple motif models can share the same TF name, so this table has 529 unique tf strings. CRE count is lower than 191,292 because this table retains the top 10% differential CREs after the 100 kb / 3-gene-per-CRE assignment used for the public network.
Data Files
| path | role | size |
|---|---|---|
network/gene_cre_tf.tsv.gz |
Core Gene–CRE–TF association table, 24,964,295 rows | ~121 MB gzipped |
metadata/conditions.tsv |
The 22 condition values in the network table |
|
metadata/columns.md |
Column-level schema and coordinate conventions | |
LICENSE |
CC BY 4.0 |
Motif-level TF binding sites are not included. They do not contain gene names and are not required to use the network table.
File Schema
See metadata/columns.md. Summary of the core table:
| column | example |
|---|---|
gene |
SORL1 |
chr |
chr20 |
start |
30502691 |
end |
30503583 |
tf |
EHF |
jaspar_id |
MA0598.4 |
condition |
SORL1KO |
The table is self-contained. Neo4j is not required to interpret a row.
Conditions
condition has 22 values, listed in metadata/conditions.tsv. Experimental design for each context is in the manuscript (Table S1).
Genome Assembly
All coordinates are GRCh38 / hg38, 1-based and closed (the start and end of a CRE are included). They are taken from the source chr:start-end interval strings without conversion to BED (0-based, half-open) coordinates.
Usage
import pandas as pd
net = pd.read_csv(
"network/gene_cre_tf.tsv.gz",
sep="\t",
dtype={"start": "int32", "end": "int32"},
)
sorl1 = net[net["gene"] == "SORL1"]
trem2 = net[net["condition"] == "TREM2KO"]
From the Hugging Face Hub:
from huggingface_hub import hf_hub_download
import pandas as pd
path = hf_hub_download(
repo_id="Thresh514/cEpiNets",
filename="network/gene_cre_tf.tsv.gz",
repo_type="dataset",
)
df = pd.read_csv(path, sep="\t")
Citation
Please cite the manuscript (DOI to be added on publication):
Fu, T.-T., Kurkela, M., Tu, J., Zhang, J., Sun, N., Farrer, L., TCW, J., and Hou, L.
Context-dependent regulatory networks connect Alzheimer's disease genetics
to microglial inflammatory states.
https://huggingface.co/datasets/Thresh514/cEpiNets
Correspondence: Lei Hou (leihou@bu.edu) and Julia TCW (juliatcw@bu.edu).
License
This dataset is released under Creative Commons Attribution 4.0 International (CC BY 4.0). See LICENSE.
- Downloads last month
- 42