【发布时间】:2020-07-08 18:52:30
【问题描述】:
该图显示了水温与时间的关系。当有激活时,温度会升高。当激活结束时,温度将开始下降(尽管有时可能会有时间延迟)。
我想计算发生事件的次数(每个蓝色圆圈代表一次激活)。有时会出现随机噪音(红色圆圈 - 表示温度随机变化,但您可以看到只有增加或减少,但不是两者兼而有之,暗示它不是一个适当的事件)。
温度每变化 0.5°C,温度记录就会更新,与时间无关。
我曾尝试使用1)温差,2)相邻数据点的温度变化梯度来识别事件开始时间戳和结束时间戳,并将其计为一个事件。但这不是很准确。
我被告知我应该只使用温差并将(增加 - 最高温度 - 降低)的模式识别为一个事件。任何想法什么是计算激活总数的合适方法?
更新1:
样本数据:
id timestamp temperature
27581 27822 2020-01-02 07:53:05.173 19.5
27582 27823 2020-01-02 07:53:05.273 20.0
27647 27888 2020-01-02 10:01:46.380 20.5
27648 27889 2020-01-02 10:01:46.480 21.0
27649 27890 2020-01-02 10:01:48.463 21.5
27650 27891 2020-01-02 10:01:48.563 22.0
27711 27952 2020-01-02 10:32:19.897 21.5
27712 27953 2020-01-02 10:32:19.997 21.0
27861 28102 2020-01-02 11:34:41.940 21.5
...
更新2:
试过了:
df['timestamp'] = pd.to_datetime(df['timestamp'])
df['Date'] = [datetime.datetime.date(d) for d in df['timestamp']]
df['Date'] = pd.to_datetime(df['Date'])
df = df[df['Date'] == '2020-01-02']
# one does not need duplicate temperature values,
# because the task is to find changing values
df2 = df.loc[df['temperature'].shift() != df['temperature']]
# ye good olde forward difference
der = np.diff(df2['temperature'])
# to have the same length as index
der = np.insert(der,len(der),np.NaN)
# make it column
df2['sig'] = np.sign(der)
# temporary array
evts = np.zeros(len(der))
# we find that points, where the signum is changing from 1 to -1, i.e. crosses zero
evts[(df2['sig'].shift() != df2['sig'])&(0 > df2['sig'])] = 1.0
# make it column for plotting
df2['events'] = evts
# preparing plot
fig,ax = plt.subplots(figsize=(20,20))
ax.xaxis_date()
ax.xaxis.set_major_locator(plticker.MaxNLocator(20))
# temperature itself
ax.plot(df2['temperature'],'-xk')
ax2=ax.twinx()
# 'events'
ax2.plot(df2['events'],'-xg')
## uncomment next two lines for plotting of signum
# ax3=ax.twinx()
# ax3.plot(df2['sig'],'-m')
# x-axis tweaking
ax.xaxis.set_major_formatter(mdates.DateFormatter('%H:%M'))
minLim = '2020-01-02 00:07:00'
maxLim = '2020-01-02 23:59:00'
plt.xlim(mdates.date2num(pd.Timestamp(minLim)),
mdates.date2num(pd.Timestamp(maxLim)))
plt.show()
并产生一个带有消息的空白图表:
/usr/local/lib/python3.6/dist-packages/ipykernel_launcher.py:31: SettingWithCopyWarning:
A value is trying to be set on a copy of a slice from a DataFrame.
Try using .loc[row_indexer,col_indexer] = value instead
See the caveats in the documentation: http://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html#returning-a-view-versus-a-copy
/usr/local/lib/python3.6/dist-packages/ipykernel_launcher.py:38: SettingWithCopyWarning:
A value is trying to be set on a copy of a slice from a DataFrame.
Try using .loc[row_indexer,col_indexer] = value instead
See the caveats in the documentation: http://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html#returning-a-view-versus-a-copy
更新3:
编写一个 for 循环来生成每天的图表:
df['timestamp'] = pd.to_datetime(df['timestamp'])
df['Date'] = df['timestamp'].dt.date
df.set_index(df['timestamp'], inplace=True)
start_date = pd.to_datetime('2020-01-01 00:00:00')
end_date = pd.to_datetime('2020-02-01 00:00:00')
df = df.loc[(df.index >= start_date) & (df.index <= end_date)]
for date in df['Date'].unique():
df_date = df[df['Date'] == date]
# one does not need duplicate temperature values,
# because the task is to find changing values
df2 = pd.DataFrame.copy(df_date.loc[df_date['temperature'].shift() != df_date['temperature']])
# ye good olde forward difference
der = np.sign(np.diff(df2['temperature']))
# to have the same length as index
der = np.insert(der,len(der),np.NaN)
# make it column
df2['sig'] = der
# temporary array
evts = np.zeros(len(der))
# we find that points, where the signum is changing from 1 to -1, i.e. crosses zero
evts[(df2['sig'].shift() != df2['sig'])&(0 > df2['sig'])] = 1.0
# make it column for plotting
df2['events'] = evts
# preparing plot
fig,ax = plt.subplots(figsize=(30,10))
ax.xaxis_date()
# df2['timestamp'] = pd.to_datetime(df2['timestamp'])
ax.xaxis.set_major_locator(plticker.MaxNLocator(20))
# temperature itself
ax.plot(df2['temperature'],'-xk')
ax2=ax.twinx()
# 'events'
g= ax2.plot(df2['events'],'-xg')
# x-axis tweaking
ax.xaxis.set_major_formatter(mdates.DateFormatter('%H:%M'))
minLim = '2020-01-02 00:07:00'
maxLim = '2020-01-02 23:59:00'
plt.xlim(mdates.date2num(pd.Timestamp(minLim)),
mdates.date2num(pd.Timestamp(maxLim)))
ax.autoscale()
plt.title(date)
print(np.count_nonzero(df2['events'][minLim:maxLim]))
plt.show(g)
图表有效,但计数无效。
更新4:
看起来有些图表(例如 2020-01-01、2020-01-04、2020-01-05)是在随机的时间片段上(可能是在周末)。有没有办法删除这些天?
【问题讨论】:
-
只是为了确保 2 看起来像是随机的温度变化,或者应该在图片中的 3 开始时完成。您还可以提供任何示例数据吗?
-
@FBruzzesi 是的,你是对的,激活或事件通常很短,持续几秒钟,但温度下降缓慢,由 2 和 3 之间的向下倾斜曲线表示。请参阅样本数据的编辑问题
-
第一个“红色”圆圈看起来与“蓝色”#2、6 和 3 相同。我是否理解正确,唯一的区别是峰值之后的斜率符号?
-
@Suthiro 是的,我想你可以这么说。另一个原因是,通过查看图表,我们注意到 red1 本身是孤立的(梯度很小),这意味着它可能是由于自然随机温度波动,而不是由于激活。另一方面,蓝色 2,3,6 处于一组事件的中间(温度通常较高),不太可能在很短的时间内出现随机的温度波动,因此我们认为它们是由实际激活引起的.对不起,这有点令人困惑..
-
您能否提供一个链接到用于制作上述图片的系列?那我可以试着分析一下。否则很难提出有用的建议。
标签: python pandas algorithm numpy time-series