【发布时间】:2019-01-17 00:13:00
【问题描述】:
所以我有这个数据框:
filename width height class xmin ymin xmax ymax
0 128782.JPG 640 512 Panel 36 385 119 510
1 128782.JPG 640 512 Panel 124 388 207 510
2 128782.JPG 640 512 Panel 210 390 294 511
3 128782.JPG 640 512 Panel 294 395 380 510
4 128782.JPG 640 512 Panel 379 398 466 511
5 128782.JPG 640 512 Panel 465 402 553 510
6 128782.JPG 640 512 P+SD 552 402 638 510
7 128782.JPG 640 512 P+SD 558 264 638 404
...
...
57170 128782.JPG 640 512 P+SD 36 242 121 383
57171 128782.JPG 640 512 HS+P+SD 36 97 122 242
57172 128782.JPG 640 512 P+SD 214 106 304 250
在名为“class”的列中包含唯一值“Panel”、“P+SD”和“HS+P+SD”。我想计算这些值有多少行,所以我尝试了这个:
print(len(split_df[split_df["class"].str.contains('Panel')]))
print(len(split_df[split_df["class"].str.contains('HS+P+SD')]))
print(len(split_df[split_df["class"].str.contains('P+SD')]))
这给了我这个输出:
56988
0
0
这是不正确的,您可以根据上面提供的 DataFrame 的 sn-p 清楚地看到,为什么 Panel 的所有内容都正确计算,而其他两个“类”名称没有计算在内?
这是 split_df.info 的输出:
RangeIndex: 57172 entries, 0 to 57171
Data columns (total 8 columns):
filename 57172 non-null object
width 57172 non-null int64
height 57172 non-null int64
class 57172 non-null object
xmin 57172 non-null int64
ymin 57172 non-null int64
xmax 57172 non-null int64
ymax 57172 non-null int64
dtypes: int64(6), object(2)
memory usage: 3.5+ MB
我终其一生都无法弄清楚哪里出了问题。任何帮助表示赞赏。
【问题讨论】:
-
我将 + 换成了反斜杠 (/)。这似乎解决了我的计数问题,为什么 + 会干扰这个?