【发布时间】:2020-04-17 15:14:00
【问题描述】:
总结:我有一个数据集,它的收集方式使得维度最初不可用。我想提取本质上是一大块未区分的数据并为其添加维度,以便对其进行查询、子集化等。这是以下问题的核心。
这是我拥有的一个 xarray 数据集:
<xarray.Dataset>
Dimensions: (chain: 1, draw: 2000, rows: 24000)
Coordinates:
* chain (chain) int64 0
* draw (draw) int64 0 1 2 3 4 5 6 7 ... 1993 1994 1995 1996 1997 1998 1999
* rows (rows) int64 0 1 2 3 4 5 6 ... 23994 23995 23996 23997 23998 23999
Data variables:
obs (chain, draw, rows) float64 4.304 3.985 4.612 ... 6.343 5.538 6.475
Attributes:
created_at: 2019-12-27T17:16:13.847972
inference_library: pymc3
inference_library_version: 3.8
这里的rows 维度对应于我需要还原到数据的多个子维度。特别是,这 24,000 行对应于来自 240 个条件的 100 个样本(这 100 个样本位于连续块中)。这些条件是gate、input、growth medium 和od 的组合。
我想得到这样的结果:
<xarray.Dataset>
Dimensions: (chain: 1, draw: 2000, gate: 1, input: 4, growth_medium: 3, sample: 100, rows: 24000)
Coordinates:
* chain (chain) int64 0
* draw (draw) int64 0 1 2 3 4 5 6 7 ... 1993 1994 1995 1996 1997 1998 1999
* rows *MultiIndex*
* gate (gate) int64 'AND'
* input (input) int64 '00', '01', '10', '11'
* growth_medium (growth_medium) 'standard', 'rich', 'slow'
* sample (sample) int64 0 1 2 3 4 5 6 7 ... 95 96 97 98 99
Data variables:
obs (chain, draw, gate, input, growth_medium, samples) float64 4.304 3.985 4.612 ... 6.343 5.538 6.475
Attributes:
created_at: 2019-12-27T17:16:13.847972
inference_library: pymc3
inference_library_version: 3.8
我有一个 pandas 数据框,它指定门、输入和生长介质的值——每一行给出一组门、输入和生长介质的值,以及一个指定位置的索引(在 rows ) 出现相应的 100 个样本集。目的是这个数据框是标注数据集的指南。
我查看了有关“重塑和重组数据”的 xarray 文档,但我不知道如何组合这些操作来完成我需要的操作。我怀疑我需要以某种方式将这些与GroupBy 结合起来,但我不明白如何。谢谢!
稍后:我有一个解决这个问题的方法,但是太恶心了,我希望有人能解释我的错误,还有什么更优雅的方法是可能的。
所以,首先,我将原始 Dataset 中的所有数据提取为原始 numpy 形式:
foo = qm.idata.posterior_predictive['obs'].squeeze('chain').values.T
foo.shape # (24000, 2000)
然后我根据需要重新塑造它:
bar = np.reshape(foo, (240, 100, 2000))
这给了我大致想要的形状:有 240 种不同的实验条件,每个都有 100 个变体,对于这些变体中的每一个,我的数据集中都有 2000 个 Monte Carlo 样本。
现在,我从熊猫DataFrame中提取了240个实验条件的信息:
import pandas as pd
# qdf is the original dataframe with the experimental conditions and some
# extraneous information in other columns
new_df = qdf[['gate', 'input', 'output', 'media', 'od_lb', 'od_ub', 'temperature']]
idx = pd.MultiIndex.from_frame(new_df)
最后,我从 numpy 数组和 pandas MultiIndex 重新组装了一个 DataArray:
xr.DataArray(bar, name='obs', dims=['regions', 'conditions', 'draws'],
coords={'regions': idx, 'conditions': range(100), 'draws': range(2000)})
如我所愿,生成的DataArray 具有这些坐标:
Coordinates:
* regions (regions) MultiIndex
- gate (regions) object 'AND' 'AND' 'AND' 'AND' ... 'AND' 'AND' 'AND'
- input (regions) object '00' '10' '10' '10' ... '01' '01' '11' '11'
- output (regions) object '0' '0' '0' '0' '0' ... '0' '0' '0' '1' '1'
- media (regions) object 'standard_media' ... 'high_osm_media_five_percent'
- od_lb (regions) float64 0.0 0.001 0.001 ... 0.0001 0.0051 0.0051
- od_ub (regions) float64 0.0001 0.0051 0.0051 2.0 ... 0.0003 2.0 2.0
- temperature (regions) int64 30 30 37 30 37 30 37 ... 37 30 37 30 37 30 37
* conditions (conditions) int64 0 1 2 3 4 5 6 7 ... 92 93 94 95 96 97 98 99
* draws (draws) int64 0 1 2 3 4 5 6 ... 1994 1995 1996 1997 1998 1999
不过,这太可怕了,而且我必须穿透xarray 抽象的所有漂亮层才能达到这一点,这似乎是错误的。尤其是因为这似乎不是科学工作流程的一个不寻常部分:获取相对原始的数据集以及需要与数据组合的元数据电子表格。那么我做错了什么?更优雅的解决方案是什么?
【问题讨论】:
标签: python pandas python-xarray