【发布时间】:2015-07-11 13:33:18
【问题描述】:
我想比较两个数据集 df1 和 df2,这样,df2$ID 中的唯一字符应作为新列添加到 df1 并为每个数据集分配 df2$Xp 值基因,如果df1的坐标与df2的坐标重叠:
df1 <- read.table(text="
Gene chr Start End
Gm12724 4 1000 1105
Zfhx2 4 1254 1369
Usp17lc 7 5004 5412
Lingo1 7 5698 5789
Sart3 7 5987 6041
Olfr978 4 1452 1564
", header=T)
df2 <- read.table(text="
ID chr Start End Xp
S8411 4 989 1258 0.312
S8411 4 1300 1800 0.144
S8411 7 5641 6874 0.136
S8413 4 1307 1360 -1.999
",header=T)
预期输出
df3 <- read.table("
Gene chr Start End S8411 S8413
Gm12724 4 1000 1105 0.312 0
Zfhx2 4 1294 1369 0.144 -1.999
Usp17lc 7 5004 5412 0 0
Lingo1 7 5698 5789 0.136 0
Sart3 7 5987 6041 0.136 0
Olfr978 4 1452 1564 0.144 0
",header=T)
【问题讨论】:
-
@akrun。我知道 findOverlaps 给出了重叠区域。从那我如何通过在 df1 中添加 unique(df2$ID) 作为新列来创建新数据集
标签: r bioinformatics bioconductor