【发布时间】:2021-05-04 10:28:28
【问题描述】:
我有这个数据集:
| Id | PrevId | NextId | Product | Process | Date |
|---|---|---|---|---|---|
| 1 | NULL | 4 | Product 1 | Process A | 2021-04-24 |
| 2 | NULL | 3 | Product 2 | Process A | 2021-04-24 |
| 3 | 2 | 5 | Product 2 | Process A | 2021-04-24 |
| 4 | 1 | 7 | Product 1 | Process B | 2021-04-26 |
| 5 | 3 | 6 | Product 2 | Process B | 2021-04-24 |
| 6 | 5 | NULL | Product 2 | Process B | 2021-04-24 |
| 7 | 4 | 9 | Product 1 | Process B | 2021-04-29 |
| 9 | 7 | 10 | Product 1 | Process A | 2021-05-01 |
| 10 | 9 | 15 | Product 1 | Process A | 2021-05-03 |
| 15 | 10 | 19 | Product 1 | Process A | 2021-05-04 |
| 19 | 15 | NULL | Product 1 | Process C | 2021-05-05 |
每个产品,我需要标记具有相同流程的连续/孤岛记录,例如:
| Id | PrevId | NextId | Product | Process | Date | Tag |
|---|---|---|---|---|---|---|
| 1 | NULL | 4 | Product 1 | Process A | 2021-04-24 | 1 |
| 4 | 1 | 7 | Product 1 | Process B | 2021-04-26 | 2 |
| 7 | 4 | 9 | Product 1 | Process B | 2021-04-29 | 2 |
| 9 | 7 | 10 | Product 1 | Process A | 2021-05-01 | 3 |
| 10 | 9 | 15 | Product 1 | Process A | 2021-05-03 | 3 |
| 15 | 10 | 19 | Product 1 | Process A | 2021-05-04 | 3 |
| 19 | 15 | NULL | Product 1 | Process C | 2021-05-05 | 4 |
一个产品要经过多个流程-es,并且可以不止一次地经过同一个流程。
我基本上需要生成 Tag 列,其背后的逻辑是将具有相同 Process 的连续记录组合在一起,但需要注意的是相同的进程可以出现在更下方,但应被视为一个新组。
我已经尝试过基本的窗口函数(ROW_NUMBER 和 DENSE_RANK),但问题是这些函数计数在分区内,而不是跨分区。 p>
【问题讨论】:
标签: sql sql-server gaps-and-islands