【问题标题】:How to change the first row to be the header in R?如何将第一行更改为 R 中的标题?
【发布时间】:2014-06-06 05:19:51
【问题描述】:

我有下表:

     X.5       X.6       X.7       X.8          X.9 X.10         X.11  X.12   X.13
17   Zip CuCurrent PaCurrent PoCurrent      Contact  Ext          Fax email Status
18  74136         0         1         0 918-491-6998    0 918-491-6659            1
19  30329         1         0         0 404-321-5711                              1
20  74136         1         0         0 918-523-2516    0 918-523-2522            1
21  80203         0         1         0 303-864-1919    0                         1
22  80120         1         0         0 345-098-8890  456                         1

如何使第一行 'zip, cucurrent, pacurrent...' 成为列标题?

谢谢,

下面是dput(dat)

structure(list(X.5 = structure(c(26L, 14L, 6L, 14L, 17L, 16L), .Label = c("", 
"1104", "1234 I don't know Ave.", "139.98", "300 Morgan St.", 
"30329", "312.95", "4101 S. 4th Street, Traff", "500 Highway 89 North", 
"644.04", "656.73", "72160", "72336-7000", "74136", "75501", 
"80120", "80203", "877.87", "Address1", "BZip", "General Svcs Admin (WPY)", 
"InvFileName2", "LDC_Org_Cost", "N/A", "NULL", "Zip"), class = "factor"), 
    X.6 = structure(c(7L, 2L, 3L, 3L, 2L, 3L), .Label = c("", 
    "0", "1", "301 7th St. SW", "800-688-6160", "Address2", "CuCurrent", 
    "Emergency", "LDC_Cost_Adj", "Mtelemetry", "N/A", "NULL", 
    "Suite 1402"), class = "factor"), X.7 = structure(c(8L, 3L, 
    2L, 2L, 3L, 2L), .Label = c("", "0", "1", "Address3", "Cucustomer", 
    "LDC_Misc_Fee", "NULL", "PaCurrent", "Room 7512"), class = "factor"), 
    X.8 = structure(c(14L, 2L, 2L, 2L, 2L, 2L), .Label = c("", 
    "0", "100.98", "237.02", "242.33", "335.04", "50.6", "City", 
    "Durham", "LDC_FinalVolume", "Leavenwoth", "Pacustomer", 
    "Petersburg", "PoCurrent", "Prescott", "Washington"), class = "factor"), 
    X.9 = structure(c(18L, 16L, 10L, 17L, 7L, 9L), .Label = c("", 
    "0", "1", "139.98", "20024", "27701", "303-864-1919", "312.95", 
    "345-098-8890", "404-321-5711", "644.04", "656.73", "66048", 
    "86313", "877.87", "918-491-6998", "918-523-2516", "Contact", 
    "LDC_FinalCost", "PoCustomer", "Zip"), class = "factor"), 
    X.10 = structure(c(14L, 2L, 1L, 2L, 2L, 9L), .Label = c("", 
    "0", "2.620194604", "2.710064788", "2.717239052", "2.766403162", 
    "202-708-4995", "3.09912854", "456", "804-504-7200", "913-682-2000", 
    "919-956-5541", "928-717-7472", "Ext", "InvoicesNeeded", 
    "LDC_UnitPrice", "NULL", "Phone"), class = "factor"), X.11 = structure(c(7L, 
    4L, 1L, 5L, 1L, 1L), .Label = c("", " ", "1067", "918-491-6659", 
    "918-523-2522", "Ext", "Fax", "InvoiceMonths", "LDC_UnitPrice_Original", 
    "NULL", "x2951"), class = "factor"), X.12 = structure(c(13L, 
    1L, 1L, 1L, 1L, 1L), .Label = c("", "0", "100.98", "202-401-3722", 
    "237.02", "242.33", "335.04", "50.6", "716- 344-3303", "804-504-7227", 
    "913- 758-4230", "919- 956-7152", "email", "Fax", "GSA", 
    "Supp_Vol"), class = "factor"), X.13 = structure(c(10L, 2L, 
    2L, 2L, 2L, 2L), .Label = c("", "1", "15", "202-497-6164", 
    "3", "804-504-7200", "Emergency", "MajorTypeId", "NULL", 
    "Status", "Supp_Vol_Adj"), class = "factor")), .Names = c("X.5", 
"X.6", "X.7", "X.8", "X.9", "X.10", "X.11", "X.12", "X.13"), row.names = 17:22, class = "data.frame")

【问题讨论】:

  • 桌子的dput() 会有所帮助。但是您可以使用colnames(dat) <- as.character(dat[1,]) 将列名和正常的 R 语法设置为“删除”第一行。
  • @RichardScriven 好点。回想起来,这看起来确实像 read.… 函数中缺少的 header=TRUE,但为什么行名从 17 开始呢?
  • @hrbrmstr as.character(dat[1,]) 返回第一行中因子水平的数字索引(作为字符)。我不知道为什么会这样。
  • 很奇怪。我确保它做了我认为它在 netflow 记录的巨大 data.table 上所做的事情。啊。但这些都不是因素(刚刚检查过)。希望我们有一个dput() 可以工作:-)
  • @RichardScriven,dat 是 csv 文件 gas 的子集。我使用此代码子集dat = gas[17:22,7:15]。我原帖中的第 17 行应该是 dat 的标题,但我不知道如何更改。

标签: r columnheader


【解决方案1】:

Shalini Baranwal's 答案是最好的,所以我会赞成。但是,未来的读者在运行该解决方案时可能会收到错误消息。我的错误是:

“setnames(x, value) 中的错误:传递了一个“列表”类型的向量。 需要输入‘字符’。”

为了解决这个问题,我修改后的解决方案是在第一步中添加一个 as.character() 包装器。完整解决方案如下:

第 1 步:将第一行复制到标题:

dat <- mtcars
names(dat) <- as.character(dat[1,])

第 2 步:删除第一行:

dat <- dat[-1,]

【讨论】:

  • 这应该作为对引用答案的评论或建议的编辑发布。虽然有用,但它并没有为原始问题提供独立的答案。
  • 我没有足够的声誉来评论任何帖子。我之前的答案已被编辑为独立的。谢谢
【解决方案2】:

最简洁的方法是为此目的设计了一个简单的函数。 你需要看门人包。

janitor::row_to_names(dat)

如果您希望第 n 行用于列名,则函数的第二个参数是要使用的行号。默认值为 1。

【讨论】:

  • 默认值对我来说似乎不是 1。可能需要row_number =1
【解决方案3】:

将数据导入R时请使用header=TRUE

【讨论】:

  • 我在导入csv时使用了header=T
【解决方案4】:

这可能是一种简单的方式:

第 1 步:将第一行复制到标题:

names(dat) <- dat[1,]

第 2 步:删除第一行:

dat <- dat[-1,]

【讨论】:

  • 这显然是最简单和最好的答案。谢谢
【解决方案5】:

如果您能够将文件中的数据重新读取到 R 中,您也可以将“skip”参数添加到 read.csv 以跳过前 16 行并使用第 17 行作为标题:

dat=read.csv("contacts.csv", skip=16, nrows=5, header=TRUE)

【讨论】:

    【解决方案6】:

    如果您不想将数据重新读入 R(看起来您不是来自 cmets),您可以执行以下操作。我必须添加一些零才能完全读取您的数据,因此请忽略这些。

    dat
    ##       V2        V3        V4        V5           V6  V7           V8    V9    V10
    ## 17   Zip CuCurrent PaCurrent PoCurrent      Contact Ext          Fax email Status
    ## 18 74136         0         1         0 918-491-6998   0 918-491-6659     0      1
    ## 19 30329         1         0         0 404-321-5711   0            0     0      1
    ## 20 74136         1         0         0 918-523-2516   0 918-523-2522     0      1
    ## 21 80203         0         1         0 303-864-1919   0            0     0      1
    ## 22 80120         1         0         0 345-098-8890 456            0     0      1
    

    首先取第一行作为列名。接下来删除第一行。通过将列转换为适当的类型来完成它。

    names(dat) <- as.matrix(dat[1, ])
    dat <- dat[-1, ]
    dat[] <- lapply(dat, function(x) type.convert(as.character(x)))
    dat
    ##     Zip CuCurrent PaCurrent PoCurrent      Contact Ext          Fax email Status
    ## 1 74136         0         1         0 918-491-6998   0 918-491-6659     0      1
    ## 2 30329         1         0         0 404-321-5711   0            0     0      1
    ## 3 74136         1         0         0 918-523-2516   0 918-523-2522     0      1
    ## 4 80203         0         1         0 303-864-1919   0            0     0      1
    ## 5 80120         1         0         0 345-098-8890 456            0     0      1
    

    【讨论】:

    • 有用的是库janitorclean_names 以避免lapply 语句。所以dat &lt;- janitor::clean_names(aum)
    【解决方案7】:

    如果您从 csv 文件中获取,请使用 read.csv 中的参数 'header'

    dat=read.csv("gas.csv", header=TRUE)
    

    如果您已经拥有数据并且不想/或无法以干净的方式获取它,您总是可以这样做

    dat=structure(list(X.5 = structure(c(26L, 14L, 6L, 14L, 17L, 16L), .Label = c("", "1104", "1234 I don't know Ave.", "139.98", "300 Morgan St.", "30329", "312.95", "4101 S. 4th Street, Traff", "500 Highway 89 North", "644.04", "656.73", "72160", "72336-7000", "74136", "75501", "80120", "80203", "877.87", "Address1", "BZip", "General Svcs Admin (WPY)", "InvFileName2", "LDC_Org_Cost", "N/A", "NULL", "Zip"), class = "factor"), X.6 = structure(c(7L, 2L, 3L, 3L, 2L, 3L), .Label = c("", "0", "1", "301 7th St. SW", "800-688-6160", "Address2", "CuCurrent", "Emergency", "LDC_Cost_Adj", "Mtelemetry", "N/A", "NULL", "Suite 1402"), class = "factor"), X.7 = structure(c(8L, 3L, 2L, 2L, 3L, 2L), .Label = c("", "0", "1", "Address3", "Cucustomer", "LDC_Misc_Fee", "NULL", "PaCurrent", "Room 7512"), class = "factor"), X.8 = structure(c(14L, 2L, 2L, 2L, 2L, 2L), .Label = c("", "0", "100.98", "237.02", "242.33", "335.04", "50.6", "City", "Durham", "LDC_FinalVolume", "Leavenwoth", "Pacustomer", "Petersburg", "PoCurrent", "Prescott", "Washington"), class = "factor"), X.9 = structure(c(18L, 16L, 10L, 17L, 7L, 9L), .Label = c("", "0", "1", "139.98", "20024", "27701", "303-864-1919", "312.95", "345-098-8890", "404-321-5711", "644.04", "656.73", "66048", "86313", "877.87", "918-491-6998", "918-523-2516", "Contact", "LDC_FinalCost", "PoCustomer", "Zip"), class = "factor"), X.10 = structure(c(14L, 2L, 1L, 2L, 2L, 9L), .Label = c("", "0", "2.620194604", "2.710064788", "2.717239052", "2.766403162", "202-708-4995", "3.09912854", "456", "804-504-7200", "913-682-2000", "919-956-5541", "928-717-7472", "Ext", "InvoicesNeeded", "LDC_UnitPrice", "NULL", "Phone"), class = "factor"), X.11 = structure(c(7L, 4L, 1L, 5L, 1L, 1L), .Label = c("", " ", "1067", "918-491-6659", "918-523-2522", "Ext", "Fax", "InvoiceMonths", "LDC_UnitPrice_Original", "NULL", "x2951"), class = "factor"), X.12 = structure(c(13L, 1L, 1L, 1L, 1L, 1L), .Label = c("", "0", "100.98", "202-401-3722", "237.02", "242.33", "335.04", "50.6", "716- 344-3303", "804-504-7227", "913- 758-4230", "919- 956-7152", "email", "Fax", "GSA", "Supp_Vol"), class = "factor"), X.13 = structure(c(10L, 2L, 2L, 2L, 2L, 2L), .Label = c("", "1", "15", "202-497-6164", "3", "804-504-7200", "Emergency", "MajorTypeId", "NULL", "Status", "Supp_Vol_Adj"), class = "factor")), .Names = c("X.5", "X.6", "X.7", "X.8", "X.9", "X.10", "X.11", "X.12", "X.13"), row.names = 17:22, class = "data.frame")
    dat2 = dat[2:6,]   
    colnames(dat2) = dat[1,] 
    dat2
    

    【讨论】:

    • 您提供的 data.frame 没有 22 行...您的问题表述不正确。既然您提供了 dput,我已经更改了答案
    猜你喜欢
    • 2021-04-25
    • 1970-01-01
    • 2011-02-13
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多