Dataset Viewer
Auto-converted to Parquet Duplicate
name
string
CID
int64
CAS
string
SMILES
string
num_atoms
int64
MW
float64
LogP
float64
TPSA
float64
HBD
int64
HBA
int64
RotBonds
int64
bee_pLD50
float64
bee_safe_label
int64
aquatic_pLC50
float64
aquatic_safe_label
int64
mammal_pLD50
float64
human_safe_label
int64
herbicide
int64
fungicide
int64
insecticide
int64
acaricide
int64
nematicide
int64
rodenticide
int64
plant_growth_regulator
int64
bactericide
int64
molluscicide
int64
ToxCast_0
0
ToxCast_0
O=[N+]([O-])c1ccc(Cl)cc1
10
157.56
2.25
43.14
0
2
1
4.402
0
4.743
0
3.174
0
0
0
0
0
0
0
0
0
0
ToxCast_1
0
ToxCast_1
C[SiH](C)O[Si](C)(C)O[Si](C)(C)O[SiH](C)C
15
282.64
2.41
27.69
0
3
6
5.66
1
5.151
1
3.222
0
0
0
0
0
0
0
0
0
0
ToxCast_2
0
ToxCast_2
CN1CCN(C(=O)C2CCCCC2)CC1
15
210.32
1.34
23.55
0
2
1
4.308
0
4.33
0
2.902
0
0
0
0
0
0
0
0
0
0
ToxCast_3
0
ToxCast_3
Nc1ccc([N+](=O)[O-])cc1
10
138.13
1.18
69.16
1
3
1
3.648
0
4.052
0
3.053
0
0
0
0
0
0
0
0
0
0
ToxCast_4
0
ToxCast_4
O=[N+]([O-])c1ccc(O)cc1
10
139.11
1.3
63.37
1
3
1
3.755
0
4.128
0
3.09
0
0
0
0
0
0
0
0
0
0
ToxCast_5
0
ToxCast_5
Cc1cc(=O)[nH]o1
7
99.09
0.28
46
1
2
0
3.324
0
3.414
0
2.783
0
0
0
0
0
0
0
0
0
0
ToxCast_6
0
ToxCast_6
COc1ccc(C(C)=O)cc1
11
150.18
1.9
26.3
0
2
2
4.364
0
4.514
0
3.069
0
0
0
0
0
0
0
0
0
0
ToxCast_7
0
ToxCast_7
CN(C)c1ccc(C=O)cc1
11
149.19
1.57
20.31
0
2
2
4.261
0
4.312
0
2.97
0
0
0
0
0
0
0
0
0
0
ToxCast_8
0
ToxCast_8
O=[N+]([O-])c1ccc(CBr)cc1
11
216.03
2.49
43.14
0
2
2
4.678
0
5.034
1
3.247
0
0
0
0
0
0
0
0
0
0
ToxCast_9
0
ToxCast_9
O=[N+]([O-])c1ccc(CCl)cc1
11
171.58
2.33
43.14
0
2
2
4.481
0
4.829
0
3.2
0
0
0
0
0
0
0
0
0
0
ToxCast_10
0
ToxCast_10
CNc1ccc([N+](=O)[O-])cc1
11
152.15
1.64
55.17
1
3
2
4.011
0
4.362
0
3.191
0
0
0
0
0
0
0
0
0
0
ToxCast_12
0
ToxCast_12
CC(C)c1ccc(C(C)C)cc1
12
162.28
3.93
0
0
0
2
5.534
1
5.766
1
3.68
0
0
0
0
0
0
0
0
0
0
ToxCast_13
0
ToxCast_13
CC(=O)c1ccc([N+](=O)[O-])cc1
12
165.15
1.8
60.21
0
3
2
4.079
0
4.491
0
3.039
0
0
0
0
0
0
0
0
0
0
ToxCast_14
0
ToxCast_14
O=C(Cl)c1ccc(C(=O)Cl)cc1
12
203.02
2.44
34.14
0
2
2
4.696
0
4.974
0
3.233
0
0
0
0
0
0
0
0
0
0
ToxCast_15
0
ToxCast_15
O=C(O)c1ccc(C(=O)O)cc1
12
166.13
1.08
74.6
2
2
2
3.64
0
4.065
0
3.225
0
0
0
0
0
0
0
0
0
0
ToxCast_16
0
ToxCast_16
CN(C)c1ccc(N(C)C)cc1
12
164.25
1.82
6.48
0
2
2
4.534
0
4.502
0
3.046
0
0
0
0
0
0
0
0
0
0
ToxCast_17
0
ToxCast_17
CCCCCCCC(OC)OC
12
174.28
2.97
18.46
0
2
8
5.679
1
5.215
1
3.39
0
0
0
0
0
0
0
0
0
0
ToxCast_18
0
ToxCast_18
O=[N+]([O-])[O-]
4
62
-0.24
66.2
0
3
0
2.818
0
3.012
0
2.428
0
0
0
0
0
0
0
0
0
0
ToxCast_19
0
ToxCast_19
CCN(CC)CCCC(C)Nc1cc(C=Cc2ccccc2Cl)nc2cc(Cl)ccc12
31
456.46
7.63
28.16
1
3
10
8.505
1
8.722
1
4.99
1
0
0
0
0
0
0
0
0
0
ToxCast_21
0
ToxCast_21
O=[N+]([O-])c1ccc([N+](=O)[O-])cc1
12
168.11
1.5
86.28
0
4
2
3.738
0
4.322
0
2.951
0
0
0
0
0
0
0
0
0
0
ToxCast_24
0
ToxCast_24
O=P(Cl)(Cl)Cl
5
153.33
2.81
17.07
0
1
0
4.861
0
5.07
1
3.343
0
0
0
0
0
0
0
0
0
0
ToxCast_26
0
ToxCast_26
ClP(Cl)(Cl)(Cl)Cl
6
208.24
4.31
0
0
0
0
5.834
1
6.106
1
3.793
0
0
0
0
0
0
0
0
0
0
ToxCast_27
0
ToxCast_27
O=S(=O)([O-])[O-]
5
96.06
-1.34
80.26
0
4
0
2.304
0
2.437
0
2.099
0
0
0
0
0
0
0
0
0
0
ToxCast_28
0
ToxCast_28
CCCCCCCCCCCl
11
176.73
4.37
0
0
0
8
6.47
1
6.061
1
3.81
0
0
0
0
0
0
0
0
0
0
ToxCast_29
0
ToxCast_29
O=[N+]([O-])c1ccc(CCO)cc1
12
167.16
1.13
63.37
1
3
3
3.758
0
4.096
0
3.039
0
0
0
0
0
0
0
0
0
0
ToxCast_30
0
ToxCast_30
CCCCCCCCCCCCCCC(=O)O
17
242.4
5.16
37.3
1
1
13
6.705
1
6.703
1
4.249
0
0
0
0
0
0
0
0
0
0
ToxCast_31
0
ToxCast_31
CCc1c2c(nc3ccc(OC(=O)N4CCC(N5CCCCC5)CC4)cc13)-c1cc3c(c(=O)n1C2)COC(=O)[C@]3(O)CC
43
586.69
4.09
114.2
1
8
4
5.866
1
6.921
1
3.927
0
0
0
0
0
0
0
0
0
0
ToxCast_32
0
ToxCast_32
CCOc1ccc([N+](=O)[O-])cc1
12
167.16
1.99
52.37
0
3
3
4.238
0
4.614
0
3.098
0
0
0
0
0
0
0
0
0
0
ToxCast_33
0
ToxCast_33
Cc1cccn2c(=O)c(-c3nnn[n-]3)cnc12
17
227.21
-0.19
87.14
0
5
1
3.138
0
3.455
0
2.444
0
0
0
0
0
0
0
0
0
0
ToxCast_34
0
ToxCast_34
CCCCCC/C=C/CCCCCCCC(=O)O
18
254.41
5.33
37.3
1
1
13
6.814
1
6.833
1
4.298
0
0
0
0
0
0
0
0
0
0
ToxCast_35
0
ToxCast_35
CCCCCCOC(=O)C(C)CC
13
186.29
3.16
26.3
0
2
7
5.733
1
5.359
1
3.447
0
0
0
0
0
0
0
0
0
0
ToxCast_36
0
ToxCast_36
O=S(=O)(O)O
5
98.08
-0.65
74.6
2
2
0
2.665
0
2.854
0
2.704
0
0
0
0
0
0
0
0
0
0
ToxCast_37
0
ToxCast_37
CCN(CC)CCN
8
116.21
0.29
29.26
1
2
4
3.517
0
3.463
0
2.786
0
0
0
0
0
0
0
0
0
0
ToxCast_38
0
ToxCast_38
CCN(CC)CCO
8
117.19
0.32
23.47
1
2
4
3.583
0
3.485
0
2.796
0
0
0
0
0
0
0
0
0
0
ToxCast_39
0
ToxCast_39
BrCc1ccccc1
8
171.04
2.58
0
0
0
1
4.95
0
4.976
0
3.274
0
0
0
0
0
0
0
0
0
0
ToxCast_41
0
ToxCast_41
C=CC1CC=CCC1
8
108.18
2.53
0
0
0
1
4.747
0
4.788
0
3.259
0
0
0
0
0
0
0
0
0
0
ToxCast_42
0
ToxCast_42
O=S(=O)([O-])Oc1ccc(C(c2ccc(OS(=O)(=O)[O-])cc2)c2ccccn2)cc1
29
435.44
1.94
145.75
0
9
7
4.902
0
5.252
1
3.082
0
0
0
0
0
0
0
0
0
0
ToxCast_43
0
ToxCast_43
CCc1ccccc1
8
106.17
2.25
0
0
0
1
4.615
0
4.615
0
3.175
0
0
0
0
0
0
0
0
0
0
ToxCast_44
0
ToxCast_44
C=Cc1ccccc1
8
104.15
2.33
0
0
0
1
4.646
0
4.658
0
3.199
0
0
0
0
0
0
0
0
0
0
ToxCast_45
0
ToxCast_45
CCCCCC(CO)CCC
11
158.28
2.98
20.23
1
1
7
5.623
1
5.181
1
3.593
0
0
0
0
0
0
0
0
0
0
ToxCast_46
0
ToxCast_46
OB(O)O
4
61.83
-2.05
60.69
3
3
0
2.048
0
1.924
0
2.484
0
0
0
0
0
0
0
0
0
0
ToxCast_48
0
ToxCast_48
C=Cc1ccncc1
8
105.14
1.72
12.89
0
1
1
4.269
0
4.298
0
3.017
0
0
0
0
0
0
0
0
0
0
ToxCast_49
0
ToxCast_49
ClCc1ccccc1
8
126.59
2.43
0
0
0
1
4.753
0
4.772
0
3.228
0
0
0
0
0
0
0
0
0
0
ToxCast_50
0
ToxCast_50
NCc1ccccc1
8
107.16
1.15
26.02
1
1
1
3.905
0
3.955
0
3.044
0
0
0
0
0
0
0
0
0
0
ToxCast_51
0
ToxCast_51
N#Cc1ccccc1
8
103.12
1.56
23.79
0
1
0
4.098
0
4.193
0
2.967
0
0
0
0
0
0
0
0
0
0
ToxCast_52
0
ToxCast_52
N#Cc1ccncc1
8
104.11
0.95
36.68
0
2
0
3.721
0
3.832
0
2.786
0
0
0
0
0
0
0
0
0
0
ToxCast_53
0
ToxCast_53
Cc1ncc(CSSCc2cnc(C)c(O)c2CO)c(CO)c1O
24
368.48
2.57
106.7
4
8
7
5.321
1
5.464
1
4.071
0
0
0
0
0
0
0
0
0
0
ToxCast_54
0
ToxCast_54
O=CC1CC=CCC1
8
110.16
1.54
17.07
0
1
1
4.166
0
4.2
0
2.962
0
0
0
0
0
0
0
0
0
0
ToxCast_55
0
ToxCast_55
OCc1ccccc1
8
108.14
1.18
20.23
1
1
1
3.971
0
3.978
0
3.054
0
0
0
0
0
0
0
0
0
0
ToxCast_56
0
ToxCast_56
CC[N+](C)(CC)CC
8
116.23
1.49
0
0
0
3
4.304
0
4.186
0
2.948
0
0
0
0
0
0
0
0
0
0
ToxCast_57
0
ToxCast_57
O=Cc1ccccc1
8
106.12
1.5
17.07
0
1
1
4.136
0
4.165
0
2.95
0
0
0
0
0
0
0
0
0
0
ToxCast_58
0
ToxCast_58
SCc1ccccc1
8
124.21
2.12
0
1
1
1
4.607
0
4.58
0
3.335
0
0
0
0
0
0
0
0
0
0
ToxCast_59
0
ToxCast_59
N#Cc1cccnc1
8
104.11
0.95
36.68
0
2
0
3.721
0
3.832
0
2.786
0
0
0
0
0
0
0
0
0
0
ToxCast_60
0
ToxCast_60
OCc1cccnc1
8
109.13
0.57
33.12
1
2
1
3.594
0
3.617
0
2.872
0
0
0
0
0
0
0
0
0
0
ToxCast_62
0
ToxCast_62
Cl/C=C\CCl
5
110.97
1.98
0
0
0
1
4.507
0
4.464
0
3.093
0
0
0
0
0
0
0
0
0
0
ToxCast_63
0
ToxCast_63
CNc1ccccc1
8
107.16
1.73
12.03
1
1
1
4.284
0
4.305
0
3.218
0
0
0
0
0
0
0
0
0
0
ToxCast_64
0
ToxCast_64
NNc1ccccc1
8
108.14
0.97
38.05
2
2
1
3.729
0
3.854
0
3.192
0
0
0
0
0
0
0
0
0
0
ToxCast_65
0
ToxCast_65
ON=C1CCCCC1
8
113.16
1.78
32.59
1
2
0
4.153
0
4.351
0
3.234
0
0
0
0
0
0
0
0
0
0
ToxCast_66
0
ToxCast_66
Clc1ccc2c(c1)CCc1cccnc1C2=C1CCNCC1
22
310.83
4.02
24.92
1
2
0
5.789
1
6.188
1
3.906
0
0
0
0
0
0
0
0
0
0
ToxCast_67
0
ToxCast_67
ONc1ccccc1
8
109.13
1.49
32.26
2
2
1
4.012
0
4.165
0
3.346
0
0
0
0
0
0
0
0
0
0
ToxCast_68
0
ToxCast_68
c1ccc(C2(c3ccccc3)CC2C2=NCCN2)cc1
20
262.36
2.99
24.39
1
2
3
5.194
1
5.452
1
3.598
0
0
0
0
0
0
0
0
0
0
ToxCast_69
0
ToxCast_69
C=Cc1ccccn1
8
105.14
1.72
12.89
0
1
1
4.269
0
4.298
0
3.017
0
0
0
0
0
0
0
0
0
0
ToxCast_70
0
ToxCast_70
N#Cc1ccccn1
8
104.11
0.95
36.68
0
2
0
3.721
0
3.832
0
2.786
0
0
0
0
0
0
0
0
0
0
ToxCast_71
0
ToxCast_71
CCNc1nc(N)nc(Cl)n1
11
173.61
0.54
76.72
2
5
2
3.399
0
3.757
0
3.062
0
0
0
0
0
0
0
0
0
0
ToxCast_72
0
ToxCast_72
CCN1CCOCC1
8
115.18
0.34
12.47
0
2
1
3.677
0
3.491
0
2.602
0
0
0
0
0
0
0
0
0
0
ToxCast_73
0
ToxCast_73
O=NN1CCCCC1
8
114.15
1.15
32.67
0
2
1
3.873
0
3.978
0
2.846
0
0
0
0
0
0
0
0
0
0
ToxCast_74
0
ToxCast_74
C1CN2CCC1CC2
8
111.19
1.1
3.24
0
1
0
4.087
0
3.939
0
2.831
0
0
0
0
0
0
0
0
0
0
ToxCast_75
0
ToxCast_75
COC(=O)c1c(Cl)nn(C)c1S(=O)(=O)NC(=O)Nc1nc(OC)cc(OC)n1
28
434.82
0.18
163.63
2
10
6
3.959
0
4.194
0
2.953
0
0
0
0
0
0
0
0
0
0
ToxCast_76
0
ToxCast_76
C=Cc1cccc(C)c1
9
118.18
2.64
0
0
0
1
4.825
0
4.878
0
3.291
0
0
0
0
0
0
0
0
0
0
ToxCast_77
0
ToxCast_77
CC(C)(c1ccccc1)c1ccc(Nc2ccc(C(C)(C)c3ccccc3)cc2)cc1
31
405.59
8.08
12.03
1
1
6
8.695
1
8.863
1
5.125
1
0
0
0
0
0
0
0
0
0
ToxCast_78
0
ToxCast_78
C[N+](C)(C)Cc1ccccc1
11
150.24
1.89
0
0
0
2
4.581
0
4.511
0
3.068
0
0
0
0
0
0
0
0
0
0
ToxCast_79
0
ToxCast_79
Oc1ccccc1-c1nnco1
12
162.15
1.44
59.15
1
4
1
3.919
0
4.271
0
3.133
0
0
0
0
0
0
0
0
0
0
ToxCast_80
0
ToxCast_80
CC(C)(O)Cc1ccccc1
11
150.22
2
20.23
1
1
2
4.461
0
4.576
0
3.3
0
0
0
0
0
0
0
0
0
0
ToxCast_81
0
ToxCast_81
O=Cc1ccccc1S(=O)(=O)[O-]
12
185.18
0.4
74.27
0
4
2
3.392
0
3.705
0
2.621
0
0
0
0
0
0
0
0
0
0
ToxCast_82
0
ToxCast_82
O=S(=O)(O)NC1CCCCC1
11
179.24
0.71
66.4
2
2
2
3.579
0
3.875
0
3.113
0
0
0
0
0
0
0
0
0
0
ToxCast_83
0
ToxCast_83
CCCC(=O)OC(C)(C)Cc1ccccc1
16
220.31
3.35
26.3
0
2
5
5.918
1
5.561
1
3.505
0
0
0
0
0
0
0
0
0
0
ToxCast_84
0
ToxCast_84
CC(=O)c1ccc(C(C)=O)cc1
12
162.19
2.09
34.14
0
2
2
4.42
0
4.661
0
3.128
0
0
0
0
0
0
0
0
0
0
ToxCast_85
0
ToxCast_85
C1N2CN3CN1CN(C2)C3
10
140.19
-1.02
12.96
0
4
0
3.134
0
2.739
0
2.194
0
0
0
0
0
0
0
0
0
0
ToxCast_86
0
ToxCast_86
C[Si]1(C)N[Si](C)(C)N[Si](C)(C)N1
12
219.51
0.87
36.09
3
3
0
4.02
0
4.073
0
3.362
0
0
0
0
0
0
0
0
0
0
ToxCast_87
0
ToxCast_87
CCOC(=O)Cn1cccc1-c1nc(-c2ccc(OC)cc2)c(-c2ccc(OC)cc2)s1
32
448.54
5.53
62.58
0
6
8
7.247
1
7.437
1
4.158
0
0
0
0
0
0
0
0
0
0
ToxCast_89
0
ToxCast_89
c1ccc(OP(Oc2ccccc2)Oc2ccccc2)cc1
22
310.29
5.45
27.69
0
3
6
7.108
1
7.046
1
4.135
0
0
0
0
0
0
0
0
0
0
ToxCast_92
0
ToxCast_92
Clc1nc(Cl)nc(Nc2ccccc2Cl)n1
16
275.53
3.58
50.7
1
4
2
5.274
1
5.834
1
3.773
0
0
0
0
0
0
0
0
0
0
ToxCast_94
0
ToxCast_94
CC(Oc1cccc(Cl)c1)C(=O)O
13
200.62
2.19
46.53
1
2
3
4.472
0
4.817
0
3.358
0
0
0
0
0
0
0
0
0
0
ToxCast_95
0
ToxCast_95
CCN(Cc1cccc(S(=O)(=O)O)c1)c1ccccc1
20
291.37
2.96
57.61
1
3
5
5.684
1
5.504
1
3.588
0
0
0
0
0
0
0
0
0
0
ToxCast_96
0
ToxCast_96
Nc1ccc(Cc2ccc(N)c(Cl)c2)cc1Cl
17
267.16
3.75
52.04
2
2
2
5.317
1
5.917
1
4.025
0
0
0
0
0
0
0
0
0
0
ToxCast_97
0
ToxCast_97
CCN(CC)C(=O)[C@]1(c2ccccc2)C[C@@H]1CN
18
246.35
1.77
46.33
1
2
5
5.115
1
4.679
0
3.231
0
0
0
0
0
0
0
0
0
0
ToxCast_98
0
ToxCast_98
Oc1cccc(Nc2ccccc2)c1
14
185.23
3.14
32.26
2
2
2
4.971
0
5.345
1
3.841
0
0
0
0
0
0
0
0
0
0
ToxCast_99
0
ToxCast_99
CN(C)c1ccc(O)c2c1C[C@H]1C[C@H]3[C@H](N(C)C)C(O)=C(C(N)=O)C(=O)[C@@]3(O)C(O)=C1C2=O
33
457.48
0.19
164.63
5
9
3
3.319
0
4.256
0
3.556
0
0
0
0
0
0
0
0
0
0
ToxCast_100
0
ToxCast_100
COC(=O)c1ccccc1S(=O)(=O)NC(=O)N(C)c1nc(C)nc(OC)n1
27
395.4
0.51
140.68
1
9
5
4.187
0
4.294
0
2.853
0
1
0
0
0
0
0
0
0
0
ToxCast_101
0
ToxCast_101
O=C(Nc1ccc(Cl)cc1)Nc1ccc(Cl)c(Cl)c1
19
315.59
5.29
41.13
2
1
2
6.24
1
6.963
1
4.487
0
0
0
0
0
0
0
0
0
0
ToxCast_102
0
ToxCast_102
CC(C)OC(=O)Nc1cccc(Cl)c1
14
213.66
3.3
38.33
1
2
2
5.075
1
5.512
1
3.689
0
0
0
0
0
0
0
0
0
0
ToxCast_104
0
ToxCast_104
CN(C)C(=O)Oc1ccc[n+](C)c1
13
181.21
0.57
33.42
0
2
1
3.796
0
3.796
0
2.671
0
0
0
0
0
0
0
0
0
0
ToxCast_105
0
ToxCast_105
O=C(Nc1cccc(Cl)c1)OCC#CCCl
16
258.1
3.13
38.33
1
2
2
5.127
1
5.524
1
3.639
0
0
0
0
0
0
0
0
0
0
ToxCast_106
0
ToxCast_106
C=C[C@]1(C)C[C@@H](OC(=O)CSC(C)(C)CNC(=O)[C@H](N)C(C)C)[C@]2(C)[C@H](C)CC[C@]3(CCC(=O)[C@H]32)[C@@H](C)[C@@H]1O
39
564.83
4.5
118.72
3
7
9
6.651
1
7.115
1
4.451
0
0
0
0
0
0
0
0
0
0
ToxCast_107
0
ToxCast_107
C=CCOc1nc(OCC=C)nc(OCC=C)n1
18
249.27
1.57
66.36
0
6
9
4.864
0
4.563
0
2.97
0
0
0
0
0
0
0
0
0
0
ToxCast_108
0
ToxCast_108
CC(C=O)=Cc1ccccc1
11
146.19
2.29
17.07
0
1
2
4.605
0
4.739
0
3.187
0
0
0
0
0
0
0
0
0
0
ToxCast_109
0
ToxCast_109
CNC(C)CC1CCCCC1
11
155.28
2.56
12.03
1
1
3
4.798
0
4.927
0
3.469
0
0
0
0
0
0
0
0
0
0
ToxCast_110
0
ToxCast_110
COC(=O)Cc1ccccc1
11
150.18
1.4
26.3
0
2
2
4.141
0
4.217
0
2.921
0
0
0
0
0
0
0
0
0
0
ToxCast_111
0
ToxCast_111
CN(C)C(=O)Nc1ccccc1
12
164.21
1.78
32.34
1
1
1
4.301
0
4.479
0
3.234
0
0
0
0
0
0
0
0
0
0
ToxCast_112
0
ToxCast_112
O=C(NC(=O)c1c(F)cccc1F)Nc1ccc(Oc2ccc(C(F)(F)F)cc2Cl)cc1F
33
488.77
6.53
67.43
2
3
4
7.073
1
8.14
1
4.859
1
0
0
0
0
0
0
0
0
0
End of preview. Expand in Data Studio

AgroBench-3D

A transparent, data-centric molecular resource for safety-aware agrochemical machine learning.

AgroBench-3D consolidates curated organic molecular structures, calculated physicochemical descriptors, non-target safety proxy targets, and agricultural-role indicators into one fixed, model-ready schema. The current release contains 11,129 unique canonical organic structures and 26 fields, with no missing cells and no duplicate canonical SMILES. It is intended to reduce the repeated data-engineering effort that otherwise separates molecular graph modeling, descriptor-based QSAR, and safety-aware agrochemical-method development.

Important release boundary. The continuous toxicity-related fields in this release are descriptor-derived proxy scores, not traceable experimental LD50/LC50 measurements. The repository includes a 3D conformer archive (conformers_3d_master.sdf), but users should verify molecule-level CSV--SDF alignment, conformer coverage, and optimization status before model training. This version is suitable for data-pipeline development, reproducible curation studies, 2D/3D methodological prototyping, and conformer-aware representation learning; it must not be used for regulatory, ecological, or experimental-toxicity claims.

Dataset Details

Dataset Description

Property Value
Repository ID Hassan2007/EcoAgro3D
Display name AgroBench-3D
Task families Tabular regression; tabular classification; molecular graph learning; multi-task learning
Records 11,129 unique organic molecules
Schema 26 fields
Primary representation Canonical SMILES, calculated descriptors, labels, and weak role annotations
3D status conformers_3d_master.sdf is included; the documented workflow uses ETKDGv3 and MMFF94s
Curator and contact Hassan Ahmed Hassan Zaki β€” hassanahmed07.e9@gmail.com
ORCID 0009-0005-0306-0898
License MIT for curator-authored release materials; see the licensing note below for upstream-source obligations

AgroBench-3D is a data resource for a practical gap in agrochemical machine learning: relevant molecular structures, calculated properties, ecological-safety concepts, and agricultural-use labels are often fragmented across incompatible sources. This release establishes a stable tabular contract so users can load the same molecular identity, descriptors, proxy endpoints, and labels without independently reconstructing the schema.

Included Files

File Status Description
agrochemical_master.csv Included (1.55 MB) Master table containing 11,129 rows and 26 columns
EXP_data_source_code.ipynb Included (29.3 kB) Curation, descriptor, proxy-target, and 3D-generation workflow
conformers_3d_master.sdf Included (38.9 MB) ETKDGv3/MMFF94s conformer archive for coordinate-aware workflows

The CSV, notebook, and SDF are published at the repository root. A future revision should add a machine-readable manifest that maps every SDF record unambiguously to a CSV row through SMILES and/or a persistent molecule identifier, together with conformer-generation status and energy information.

Data Sources

The curation workflow combines a verified local agrochemical master set with structural records drawn from EPA ToxCast, Tox21, ClinTox, SIDER, BBBP, BACE, and HIV source tables. The non-agrochemical sources expand structural diversity; their inclusion must not be interpreted as evidence of agricultural efficacy or measured ecological safety for every included compound. Source snapshots, source-specific licenses, acquisition dates, and record-level assay provenance are not bundled with the present release and are required for a future experimentally grounded benchmark.

Intended Uses

Direct Use

Appropriate uses of this release include:

  • Building and validating CSV, RDKit, PyTorch Geometric, DGL, or DeepChem data loaders for agrochemical molecules.
  • Converting canonical SMILES to 2D molecular graphs for representation-learning and multi-task-learning prototypes.
  • Reproducing descriptor distributions, quality-control checks, and the published proxy-target construction.
  • Developing scaffold-splitting, leakage-detection, task-masking, calibration, and class-imbalance workflows.
  • Prototyping ranking and generative-model interfaces using proxy objectives, provided all results are described as methodological rather than toxicological findings.
  • Loading the supplied SDF for 3D graph neural networks, conformer-aware representation learning, and coordinate-conditioned generative-model engineering after verifying CSV--SDF record alignment.
  • Auditing, extending, and replacing the current proxy targets with experimentally traceable endpoints.

Recommended Evaluation Practice

The CSV does not contain predefined train/validation/test splits. For predictive experiments, create and publish fixed scaffold-disjoint splits before comparing models. When source metadata are available, supplement them with source-disjoint or time-disjoint external evaluation. Use balanced accuracy, AUROC, AUPRC, calibration, and confidence intervals as appropriate; raw accuracy is insufficient for the strongly imbalanced mammalian proxy-label task.

Out-of-Scope Use

This release must not be used to:

  • Make regulatory decisions, pesticide-registration decisions, environmental-risk assessments, or human-health claims.
  • Treat bee_pLD50, aquatic_pLC50, or mammal_pLD50 as experimentally measured molar toxicity values.
  • Represent *_safe_label fields as validated safety determinations or thresholded assay outcomes.
  • Claim that a molecule is safe for honeybees, aquatic organisms, mammals, food exposure, or any non-target species.
  • Present a single ETKDG/MMFF conformer as a complete conformational ensemble, a quantum-mechanically validated geometry, or a binding pose.
  • Treat the sparse agricultural-role fields as complete ground-truth use classifications.

Dataset Structure

Unit of Observation

Each row represents one deduplicated, standardized organic molecular structure. Deduplication is performed on canonical SMILES after largest-fragment selection.

Field Reference

Group Fields Description
Identity and structure name, CID, CAS, SMILES Compound identifiers and canonical molecular structure. CID and CAS should be independently checked against source registries for identity-critical applications.
Structural and physicochemical descriptors num_atoms, MW, LogP, TPSA, HBD, HBA, RotBonds RDKit-derived heavy-atom count, molecular weight, calculated lipophilicity, topological polar surface area, hydrogen-bond donor/acceptor counts, and rotatable-bond count.
Continuous proxy scores bee_pLD50, aquatic_pLC50, mammal_pLD50 Descriptor-derived continuous scores. Despite their names, these are not assay-derived pLD50/pLC50 measurements in the current release.
Binary proxy labels bee_safe_label, aquatic_safe_label, human_safe_label Binary values obtained by thresholding the corresponding proxy scores. In the supplied code, 1 denotes a score above its cutoff; field names should not be interpreted as validated safety annotations.
Agricultural-role indicators herbicide, fungicide, insecticide, acaricide, nematicide, rodenticide, plant_growth_regulator, bactericide, molluscicide Non-exclusive weak labels assigned using keywords and limited SMARTS patterns.

Descriptive Statistics

Property Mean Median SD Minimum Maximum
Molecular weight (Da) 304.26 284.34 148.37 60.01 796.02
LogP 2.37 2.49 2.25 -13.20 14.57
TPSA (Γ…Β²) 65.04 57.36 47.54 0.00 399.71

Binary-Label Composition

Field Label 0 Label 1 Interpretation for this release
bee_safe_label 5,632 (50.607%) 5,497 (49.393%) Near-balanced proxy task
aquatic_safe_label 4,908 (44.101%) 6,221 (55.899%) Near-balanced proxy task
human_safe_label 10,639 (95.597%) 490 (4.403%) Strongly imbalanced proxy task

Agricultural-Role Coverage

The current weak-label coverage is sparse: herbicide has 40 positive records, fungicide has 94, and insecticide has 49. The other six role fields have no positive records, and 10,946 molecules have no assigned agricultural role. Use task masking, weak-supervision methods, or a curated subset; do not report all nine fields as fully supervised classification tasks.

Dataset Creation

Curation Rationale

The resource was created to make agrochemical molecular data easier to use as shared infrastructure. The design goal is a fixed schema that joins structure, molecular properties, non-target proxy objectives, and agricultural-role indicators, reducing the effort required to begin reproducible cheminformatics and machine-learning studies.

Data Collection and Processing

The supplied implementation performs the following operations:

  1. Parses input SMILES with RDKit.
  2. Splits disconnected fragments and retains the largest heavy-atom organic fragment.
  3. Removes missing or unparsable strings, strings containing a vertical-bar delimiter, compounds outside 4--65 heavy atoms, and elements outside C, H, N, O, S, P, F, Cl, Br, I, B, and Si.
  4. Restricts molecular weight to 60--800 Da.
  5. Serializes the retained structures as canonical SMILES and removes duplicate canonical structures.
  6. Calculates molecular weight, LogP, TPSA, HBD, HBA, and rotatable-bond counts.
  7. Assigns agricultural-role indicators using name keywords and limited SMARTS patterns.
  8. Creates the current continuous and binary targets from explicit descriptor-based formulas.
  9. Specifies a one-conformer ETKDGv3 embedding workflow with random seed 42 and MMFF94s minimization for up to 150 iterations.

The documented specification proposes a Tice-inspired LogP interval of -2 to 8; the supplied code calculates LogP but does not enforce this filter. The observed range is -13.20 to 14.57. Users reproducing a strict chemical-space release should implement the filter explicitly and publish a new, versioned artifact.

Current Proxy-Target Construction

Let M denote molecular weight, L LogP, A TPSA, R rotatable bonds, and D hydrogen-bond donors. The code computes the intermediate score:

h=0.45L+M350βˆ’A120.h = 0.45L + \frac{M}{350} - \frac{A}{120}.

It then creates the clipped continuous proxy scores:

sbee=clip⁑(3.5+h+0.5I(R>4)βˆ’0.2,β€…β€Š0.5,β€…β€Š9.5),s_{\mathrm{bee}} = \operatorname{clip}\left(3.5 + h + 0.5\mathbb{I}(R > 4) - 0.2,\; 0.5,\; 9.5\right),

saquatic=clip⁑(3.0+0.6L+M400,β€…β€Š0.5,β€…β€Š10.0),s_{\mathrm{aquatic}} = \operatorname{clip}\left(3.0 + 0.6L + \frac{M}{400},\; 0.5,\; 10.0\right),

smammal=clip⁑(2.5+0.3L+0.2D,β€…β€Š0.5,β€…β€Š8.5).s_{\mathrm{mammal}} = \operatorname{clip}\left(2.5 + 0.3L + 0.2D,\; 0.5,\; 8.5\right).

The binary fields are 1 when the respective score exceeds 5.0, 5.0, and 4.5. This explicit construction is useful for reproducibility and software testing, but it creates a direct leakage pathway: a model supplied with the same descriptors can recover the generating relationships. The current proxy fields are therefore not an independent predictive-toxicology benchmark.

3D-Conformer Procedure and Release Status

The repository includes conformers_3d_master.sdf (38.9 MB). The code specifies one ETKDGv3 conformer per standardized molecule, with explicit hydrogens, a fixed random seed of 42, and MMFF94s minimization for up to 150 iterations. The archive supports coordinate-aware engineering experiments. Nevertheless, a single force-field-minimized conformer is not a conformational ensemble, a quantum-mechanically validated geometry, or a protein-bound pose. Because a separate conformer manifest is not yet included, users should verify CSV--SDF identity mapping, conformer coverage, embedding failures, and optimization status before training or reporting 3D-model results.

Data Quality Checks

The released master table was audited for:

  • Completeness: 0 null cells across 11,129 rows and 26 columns.
  • Canonical-structure uniqueness: 0 duplicate SMILES entries after curation.
  • Structural bounds: heavy-atom count from 4 to 57 and molecular weight from 60.01 to 796.02 Da.
  • Label prevalence: reported above for every binary proxy task.
  • Target dependence: the continuous proxy scores are strongly correlated because they share calculated descriptor inputs; this is expected and documented, not evidence of independent biological mechanisms.

Bias, Risks, and Limitations

Scientific and Technical Limitations

  • No assay provenance for current toxicity fields. The score names resemble standard toxicological endpoints, but the current values are generated from descriptors and are not measured pLD50/pLC50 observations.
  • Descriptor leakage. LogP, molecular weight, TPSA, HBD, and rotatable bonds directly enter the proxy formulas. Reported performance against these labels can overstate a model's ability to learn toxicological mechanisms.
  • Sparse role annotations. Only three of nine agricultural-role fields currently contain positive examples; these are weak labels rather than a complete functional ontology.
  • Single-conformer and mapping limitations. A 3D SDF is included, but a single ETKDG/MMFF conformer does not represent conformational diversity or a bound pose; a machine-readable CSV--SDF manifest and conformer quality report are still needed.
  • Unenforced advertised LogP filter. Although the intended interval is -2 to 8, the source code does not apply it.
  • Chemical identity scope. Zero duplicate canonical SMILES does not verify registry identity, stereochemical completeness, commercial availability, pesticide registration, biological activity, or environmental fate.

Responsible-Use Recommendations

Use this version as a transparent engineering resource. Clearly state that outcomes concern proxy-label prediction, data curation, conformer-aware method development, or other methodological evaluation. For ecological, toxicological, or decision-making use, first construct a new release with assay-level provenance, species and life-stage context, exposure route, endpoint definition, unit conversion, censoring status, quality flags, expert-curated agricultural roles, an audited CSV--SDF manifest and conformer quality report, and external validation.

Licensing and Redistribution Note

The mit metadata is intended for curator-authored documentation, code, and release materials. Before publishing a redistributed aggregate under this license, verify that each upstream data source permits redistribution and that its attribution, database-right, and license conditions are satisfied. If that audit identifies restrictions, change the repository license metadata and provide a source-specific licensing manifest before release.

Citation

If you use AgroBench-3D, cite the dataset release. The citation below uses @misc rather than @article because the Hugging Face release is a dataset artifact; replace the version field and add a DOI if a versioned archival release is minted.

BibTeX

@misc{zaki2026agrobench3d,
  author       = {Zaki, Hassan Ahmed Hassan},
  title        = {{AgroBench-3D}: A Data-Centric Resource for Safety-Aware Agrochemical Machine Learning},
  year         = {2026},
  month        = aug,
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/Hassan2007/EcoAgro3D}},
  note         = {Hugging Face dataset repository}
}

APA 7th edition

Zaki, H. A. H. (2026). AgroBench-3D: A data-centric resource for safety-aware agrochemical machine learning [Data set]. Hugging Face. https://huggingface.co/datasets/Hassan2007/EcoAgro3D

If a manuscript is formally published, cite both the dataset release and the paper. Do not replace the dataset citation with an article citation unless the referenced article has actually been published.

Dataset Card Authors

Hassan Ahmed Hassan Zaki β€” ORCID 0009-0005-0306-0898

Dataset Card Contact

For questions, corrections, provenance contributions, or release updates, contact hassanahmed07.e9@gmail.com.

Downloads last month
57