【发布时间】:2013-02-12 12:02:22
【问题描述】:
之前的相关问题:
Select a random entry from a group after grouping by a value (not column)?
我当前的查询如下所示:
WITH
points AS (
SELECT unnest(array_of_points) AS p
),
gtps AS (
SELECT DISTINCT ON(points.p)
points.p, m.groundtruth
FROM measurement m, points
WHERE st_distance(m.groundtruth, points.p) < distance
ORDER BY points.p, RANDOM()
)
SELECT DISTINCT ON(gtps.p, gtps.groundtruth, m.anchor_id)
m.id, m.anchor_id, gtps.groundtruth, gtps.p
FROM measurement m, gtps
ORDER BY gtps.p, gtps.groundtruth, m.anchor_id, RANDOM()
语义:
-
有两个输入值:
- 第 4 行:点数组
array_of_points - 第 12 行:双精度数:
distance
- 第 4 行:点数组
-
第一段(第 1-6 行):
- 从点数组创建一个表以用于...
-
第二段(第 8-14 行):
- 对于
points表中的每个点:从measurement表中获取一个随机(!)groundtruth点,其距离distance - 将这些元组保存在
gtps表中
- 对于
-
第三段(第 16-19 行):
- 对于
gtps表中的每个groundtruth值:获取所有anchor_id值和... - 如果
anchor_id值不唯一:则随机选择一个
- 对于
输出:
id、anchor_id、groundtruth、p(来自array_of_points的输入值)
示例表:
id | anchor_id | groundtruth | data
-----------------------------------
1 | 1 | POINT(1 4) | ...
2 | 3 | POINT(1 4) | ...
3 | 8 | POINT(1 4) | ...
4 | 6 | POINT(1 4) | ...
-----------------------------------
5 | 2 | POINT(3 2) | ...
6 | 4 | POINT(3 2) | ...
-----------------------------------
7 | 1 | POINT(4 3) | ...
8 | 1 | POINT(4 3) | ...
9 | 6 | POINT(4 3) | ...
10 | 7 | POINT(4 3) | ...
11 | 3 | POINT(4 3) | ...
-----------------------------------
12 | 1 | POINT(6 2) | ...
13 | 5 | POINT(6 2) | ...
示例结果:
id | anchor_id | groundtruth | p
-----------------------------------------
1 | 1 | POINT(1 4) | POINT(1 0)
2 | 3 | POINT(1 4) | POINT(1 0)
4 | 6 | POINT(1 4) | POINT(1 0)
3 | 8 | POINT(1 4) | POINT(1 0)
5 | 2 | POINT(3 2) | POINT(2 2)
6 | 4 | POINT(3 2) | POINT(2 2)
1 | 1 | POINT(1 4) | POINT(4 8)
2 | 3 | POINT(1 4) | POINT(4 8)
4 | 6 | POINT(1 4) | POINT(4 8)
3 | 8 | POINT(1 4) | POINT(4 8)
12 | 1 | POINT(6 2) | POINT(7 3)
13 | 5 | POINT(6 2) | POINT(7 3)
1 | 1 | POINT(4 3) | POINT(9 1)
11 | 3 | POINT(4 3) | POINT(9 1)
9 | 6 | POINT(4 3) | POINT(9 1)
10 | 7 | POINT(4 3) | POINT(9 1)
如你所见:
- 每个输入值可以有多个相等的
groundtruth值。 - 如果输入值有多个
groundtruth值,则它们必须全部相等。 - 每个 groundtruth-inputPoint-tuple 都与该 groundtruth 的每个可能的
anchor_id连接。 - 两个不同的输入值可以有相同的对应
groundtruth值。 - 两个不同的groundtruth-inputPoint-tuples 可以有相同的
anchor_id - 两个相同的 groundtruth-inputPoint-tuples 必须有不同的
anchor_ids
基准测试(针对两个输入值):
- 第 1-6 行:16 毫秒
- 第 8-14 行:48 毫秒
- 第 16-19 行:600 毫秒
详细解释:
Unique (cost=11119.32..11348.33 rows=18 width=72)
Output: m.id, m.anchor_id, gtps.groundtruth, gtps.p, (random())
CTE points
-> Result (cost=0.00..0.01 rows=1 width=0)
Output: unnest('{0101000000EE7C3F355EF24F4019390B7BDA011940:01010000003480B74082FA44402CD49AE61D173C40}'::geometry[])
CTE gtps
-> Unique (cost=7659.95..7698.12 rows=1 width=160)
Output: points.p, m.groundtruth, (random())
-> Sort (cost=7659.95..7679.04 rows=7634 width=160)
Output: points.p, m.groundtruth, (random())
Sort Key: points.p, (random())
-> Nested Loop (cost=0.00..6565.63 rows=7634 width=160)
Output: points.p, m.groundtruth, random()
Join Filter: (st_distance(m.groundtruth, points.p) < m.distance)
-> CTE Scan on points (cost=0.00..0.02 rows=1 width=32)
Output: points.p
-> Seq Scan on public.measurement m (cost=0.00..535.01 rows=22901 width=132)
Output: m.id, m.anchor_id, m.tag_node_id, m.experiment_id, m.run_id, m.anchor_node_id, m.groundtruth, m.distance, m.distance_error, m.distance_truth, m."timestamp"
-> Sort (cost=3421.18..3478.43 rows=22901 width=72)
Output: m.id, m.anchor_id, gtps.groundtruth, gtps.p, (random())
Sort Key: gtps.p, gtps.groundtruth, m.anchor_id, (random())
-> Nested Loop (cost=0.00..821.29 rows=22901 width=72)
Output: m.id, m.anchor_id, gtps.groundtruth, gtps.p, random()
-> CTE Scan on gtps (cost=0.00..0.02 rows=1 width=64)
Output: gtps.p, gtps.groundtruth
-> Seq Scan on public.measurement m (cost=0.00..535.01 rows=22901 width=8)
Output: m.id, m.anchor_id, m.tag_node_id, m.experiment_id, m.run_id, m.anchor_node_id, m.groundtruth, m.distance, m.distance_error, m.distance_truth, m."timestamp"
解释分析:
Unique (cost=11119.32..11348.33 rows=18 width=72) (actual time=548.991..657.992 rows=36 loops=1)
CTE points
-> Result (cost=0.00..0.01 rows=1 width=0) (actual time=0.004..0.011 rows=2 loops=1)
CTE gtps
-> Unique (cost=7659.95..7698.12 rows=1 width=160) (actual time=133.416..146.745 rows=2 loops=1)
-> Sort (cost=7659.95..7679.04 rows=7634 width=160) (actual time=133.415..142.255 rows=15683 loops=1)
Sort Key: points.p, (random())
Sort Method: external merge Disk: 1248kB
-> Nested Loop (cost=0.00..6565.63 rows=7634 width=160) (actual time=0.045..46.670 rows=15683 loops=1)
Join Filter: (st_distance(m.groundtruth, points.p) < m.distance)
-> CTE Scan on points (cost=0.00..0.02 rows=1 width=32) (actual time=0.007..0.020 rows=2 loops=1)
-> Seq Scan on measurement m (cost=0.00..535.01 rows=22901 width=132) (actual time=0.013..3.902 rows=22901 loops=2)
-> Sort (cost=3421.18..3478.43 rows=22901 width=72) (actual time=548.989..631.323 rows=45802 loops=1)
Sort Key: gtps.p, gtps.groundtruth, m.anchor_id, (random())"
Sort Method: external merge Disk: 4008kB
-> Nested Loop (cost=0.00..821.29 rows=22901 width=72) (actual time=133.449..166.294 rows=45802 loops=1)
-> CTE Scan on gtps (cost=0.00..0.02 rows=1 width=64) (actual time=133.420..146.753 rows=2 loops=1)
-> Seq Scan on measurement m (cost=0.00..535.01 rows=22901 width=8) (actual time=0.014..4.409 rows=22901 loops=2)
Total runtime: 834.626 ms
实时运行时,它应该运行大约 100-1000 个输入值。所以现在需要 35 到 350 秒,这太长了。
我已经尝试删除 RANDOM() 函数。这将运行时间(对于 2 个输入值)从大约 670 毫秒减少到大约 530 毫秒。所以这不是目前的主要影响。
如果这样更容易/更快,也可以运行 2 或 3 个单独的查询并在软件中执行某些部分(它在 Ruby on Rails 服务器上运行)。比如随机选择?!
正在进行的工作:
SELECT
m.groundtruth, ps.p, ARRAY_AGG(m.anchor_id), ARRAY_AGG(m.id)
FROM
measurement m
JOIN
(SELECT unnest(point_array) AS p) AS ps
ON ST_DWithin(ps.p, m.groundtruth, distance)
GROUP BY groundtruth, ps.p
使用此查询非常快(15ms),但缺少很多:
- 我只需要为每个
ps.p随机一行 - 这两个数组属于彼此。意思是:里面物品的顺序很重要!
- 这两个数组需要过滤(随机):
对于数组中出现多次的每个anchor_id:保留一个随机的并删除所有其他的。这也意味着从id-array 中为每个删除的anchor_id删除相应的id
如果anchor_id 和id 可以存储在一个元组数组中,那就太好了。例如:{[4,1],[6,3],[4,2],[8,5],[4,4]}(约束:每个元组都是唯一的,每个 id(== 示例中的第二个值)都是唯一的,anchor_ids 不是唯一的)。此示例显示没有仍必须应用的过滤器的查询。应用过滤器后,它看起来像这样{[6,3],[4,4],[8,5]}。
进行中的工作二:
SELECT DISTINCT ON (ps.p)
m.groundtruth, ps.p, ARRAY_AGG(m.anchor_id), ARRAY_AGG(m.id)
FROM
measurement m
JOIN
(SELECT unnest(point_array) AS p) AS ps
ON ST_DWithin(ps.p, m.groundtruth, distance)
GROUP BY ps.p, m.groundtruth
ORDER BY ps.p, RANDOM()
这现在给出了非常好的结果并且仍然非常快:16ms
只剩下一件事要做:
-
ARRAY_AGG(m.anchor_id)已经随机化,但是: - 它包含很多重复的条目,所以:
- 我想在上面使用 DISTINCT 之类的东西,但是:
- 它必须与
ARRAY_AGG(m.id)同步。这意味着:
如果 DISTINCT 命令保留anchor_id数组的索引 1、4 和 7,那么它还必须保留id数组的索引 1、4 和 7(当然删除所有其他索引)
【问题讨论】:
-
请运行
explain analyze以查看实际花费的时间。 -
完成。输出低于
explain verbose -
始终明确指定您的连接,即使
CROSS JOINs - 但是,如果您可以预先计算最大距离的坐标边界(包括points)。外部查询 - 可能 - 在它使用交叉连接的方式上具有误导性 - 为什么它不是基于先前收集的点?我不确定使用RANDOM()对您的结果进行排序是否真的在做您想做的事,但我不确定是否不可取... -
关于
RANDOM():输出正是我想要的。引用:“外部查询在使用交叉连接的方式上可能具有误导性——为什么它不是基于先前收集的点?”你是什么意思?我只是不知道如何将所有事情都放在一个查询中,所以我决定按照我希望它们发生的顺序写下这些事情...... -
我已经开始了另一个问题,因为这个问题现在非常具体地涉及到 postgres 和 array_agg:stackoverflow.com/questions/15102630/…
标签: sql postgresql query-optimization aggregate-functions postgis