【问题标题】:C# OPEN XML: empty cells are getting skipped while getting data from EXCEL to DATATABLEC# OPEN XML:将数据从 EXCEL 获取到 DATATABLE 时会跳过空单元格
【发布时间】:2016-07-06 03:07:44
【问题描述】:

任务

excel导入数据到DataTable

问题

不包含任何数据的单元格将被跳过,并且该行中包含数据的下一个单元格将用作空列的值。 例如

A1 为空 A2 的值为 Tom 然后在导入数据时 A1 获取 A2 的值>A2 仍然是空的

为了清楚起见,我在下面提供了一些屏幕截图

这是excel数据

这是从excel导入数据后的DataTable

代码

public class ImportExcelOpenXml
{
    public static DataTable Fill_dataTable(string fileName)
    {
        DataTable dt = new DataTable();

        using (SpreadsheetDocument spreadSheetDocument = SpreadsheetDocument.Open(fileName, false))
        {

            WorkbookPart workbookPart = spreadSheetDocument.WorkbookPart;
            IEnumerable<Sheet> sheets = spreadSheetDocument.WorkbookPart.Workbook.GetFirstChild<Sheets>().Elements<Sheet>();
            string relationshipId = sheets.First().Id.Value;
            WorksheetPart worksheetPart = (WorksheetPart)spreadSheetDocument.WorkbookPart.GetPartById(relationshipId);
            Worksheet workSheet = worksheetPart.Worksheet;
            SheetData sheetData = workSheet.GetFirstChild<SheetData>();
            IEnumerable<Row> rows = sheetData.Descendants<Row>();

            foreach (Cell cell in rows.ElementAt(0))
            {
                dt.Columns.Add(GetCellValue(spreadSheetDocument, cell));
            }

            foreach (Row row in rows) //this will also include your header row...
            {
                DataRow tempRow = dt.NewRow();

                for (int i = 0; i < row.Descendants<Cell>().Count(); i++)
                {
                    tempRow[i] = GetCellValue(spreadSheetDocument, row.Descendants<Cell>().ElementAt(i));
                }

                dt.Rows.Add(tempRow);
            }

        }

        dt.Rows.RemoveAt(0); //...so i'm taking it out here.

        return dt;
    }


    public static string GetCellValue(SpreadsheetDocument document, Cell cell)
    {
        SharedStringTablePart stringTablePart = document.WorkbookPart.SharedStringTablePart;
        string value = cell.CellValue.InnerXml;

        if (cell.DataType != null && cell.DataType.Value == CellValues.SharedString)
        {
            return stringTablePart.SharedStringTable.ChildElements[Int32.Parse(value)].InnerText;
        }
        else
        {
            return value;
        }
    }
}

我的想法

我觉得有问题

public IEnumerable&lt;T&gt; Descendants&lt;T&gt;() where T : OpenXmlElement;

如果我想要使用 Descendants 的列数

IEnumerable<Row> rows = sheetData.Descendants<<Row>();
int colCnt = rows.ElementAt(0).Count();

如果我使用 Descendants 获取行数

IEnumerable<Row> rows = sheetData.Descendants<<Row>();
int rowCnt = rows.Count();`

在这两种情况下,Descendants 都会跳过空单元格

有没有Descendants的替代品。

非常感谢您的建议

PS:我也想过通过使用像 A1, A2 这样的列名来获取单元格的值,但为了做到这一点,我必须获得列和行的确切计数,这不是可以通过使用Descendants 函数来实现。

【问题讨论】:

  • 空单元格没有 e Cell 元素,因此您找不到它们。
  • @AlexanderDerck 那么如何解决这个问题呢?
  • 使用 EPPlus 库会更容易(它使用 open xml sdk),参见示例here
  • 您还可以要求单元格始终包含一个值。如果没有标记,则默认值为零。

标签: c# datatable openxml openxml-sdk spreadsheetml


【解决方案1】:

如果一行的所有单元格中都有数据,那么一切正常。当您连续有一个空单元格时,事情就会变得混乱。

为什么会发生这种情况

这是因为在下面的代码中:

row.Descendants<Cell>().Count()

Count()非空 填充单元格(不是所有列)的数量。因此,当您将 row.Descendants&lt;Cell&gt;().ElementAt(i) 作为参数传递给 GetCellValue 方法时:

GetCellValue(spreadSheetDocument, row.Descendants<Cell>().ElementAt(i));

然后它将找到下一个非空填充单元格的内容(不一定是该列索引中的内容,i),例如如果第一列是空的,我们调用ElementAt(1),它会返回第二列中的值,整个逻辑就会混乱。

解决方案 - 我们需要处理空单元格的出现:本质上我们需要找出单元格的原始列索引,以防它之前有空单元格。因此,您需要替换您的 for 循环代码,如下所示:

for (int i = 0; i < row.Descendants<Cell>().Count(); i++)
{
      tempRow[i] = GetCellValue(spreadSheetDocument, row.Descendants<Cell>().ElementAt(i));
}

for (int i = 0; i < row.Descendants<Cell>().Count(); i++)
{
    Cell cell = row.Descendants<Cell>().ElementAt(i);
    int actualCellIndex = CellReferenceToIndex(cell);
    tempRow[actualCellIndex] = GetCellValue(spreadSheetDocument, cell);
}

并在您的代码中添加以下方法,该方法用于上述修改后的代码 sn-p 以获得任何单元格的原始/正确列索引:

private static int CellReferenceToIndex(Cell cell)
{
    int index = 0;
    string reference = cell.CellReference.ToString().ToUpper();
    foreach (char ch in reference)
    {
        if (Char.IsLetter(ch))
        {
            int value = (int)ch - (int)'A';
            index = (index == 0) ? value : ((index + 1) * 26) + value;
        }
        else
        {
            return index;
        }
    }
    return index;
}

【讨论】:

  • 我认为 CellReferenceToIndex 方法不适用于超过 az 到 aa,ab,... 的 excel,当 z col 再次超过时,它会从 0 返回索引 .... 所以如果你有很多数字的 excel无法使用的 cols
  • 谢谢你拯救了我的一天
  • 伟大的@RouzbehZarandi +1
【解决方案2】:
public void Read2007Xlsx()
        {
            try
            {
                DataTable dt = new DataTable();
                using (SpreadsheetDocument spreadSheetDocument = SpreadsheetDocument.Open(@"D:\File.xlsx", false))
                {
                    WorkbookPart workbookPart = spreadSheetDocument.WorkbookPart;
                    IEnumerable<Sheet> sheets = spreadSheetDocument.WorkbookPart.Workbook.GetFirstChild<Sheets>().Elements<Sheet>();
                    string relationshipId = sheets.First().Id.Value;
                    WorksheetPart worksheetPart = (WorksheetPart)spreadSheetDocument.WorkbookPart.GetPartById(relationshipId);
                    Worksheet workSheet = worksheetPart.Worksheet;
                    SheetData sheetData = workSheet.GetFirstChild<SheetData>();
                    IEnumerable<Row> rows = sheetData.Descendants<Row>();
                    foreach (Cell cell in rows.ElementAt(0))
                    {
                        dt.Columns.Add(GetCellValue(spreadSheetDocument, cell));
                    }
                    foreach (Row row in rows) //this will also include your header row...
                    {
                        DataRow tempRow = dt.NewRow();
                        int columnIndex = 0;
                        foreach (Cell cell in row.Descendants<Cell>())
                        {
                            // Gets the column index of the cell with data
                            int cellColumnIndex = (int)GetColumnIndexFromName(GetColumnName(cell.CellReference));
                            cellColumnIndex--; //zero based index
                            if (columnIndex < cellColumnIndex)
                            {
                                do
                                {
                                    tempRow[columnIndex] = ""; //Insert blank data here;
                                    columnIndex++;
                                }
                                while (columnIndex < cellColumnIndex);
                            }
                            tempRow[columnIndex] = GetCellValue(spreadSheetDocument, cell);

                            columnIndex++;
                        }
                        dt.Rows.Add(tempRow);
                    }
                }
                dt.Rows.RemoveAt(0); //...so i'm taking it out here.
            }
            catch (Exception ex)
            {
            }
        }
        /// <summary>
        /// Given a cell name, parses the specified cell to get the column name.
        /// </summary>
        /// <param name="cellReference">Address of the cell (ie. B2)</param>
        /// <returns>Column Name (ie. B)</returns>
        public static string GetColumnName(string cellReference)
        {
            // Create a regular expression to match the column name portion of the cell name.
            Regex regex = new Regex("[A-Za-z]+");
            Match match = regex.Match(cellReference);
            return match.Value;
        }
        /// <summary>
        /// Given just the column name (no row index), it will return the zero based column index.
        /// Note: This method will only handle columns with a length of up to two (ie. A to Z and AA to ZZ). 
        /// A length of three can be implemented when needed.
        /// </summary>
        /// <param name="columnName">Column Name (ie. A or AB)</param>
        /// <returns>Zero based index if the conversion was successful; otherwise null</returns>
        public static int? GetColumnIndexFromName(string columnName)
        {

            //return columnIndex;
            string name = columnName;
            int number = 0;
            int pow = 1;
            for (int i = name.Length - 1; i >= 0; i--)
            {
                number += (name[i] - 'A' + 1) * pow;
                pow *= 26;
            }
            return number;
        }
        public static string GetCellValue(SpreadsheetDocument document, Cell cell)
        {
            SharedStringTablePart stringTablePart = document.WorkbookPart.SharedStringTablePart;
            if (cell.CellValue ==null)
            {
            return "";
            }
            string value = cell.CellValue.InnerXml;
            if (cell.DataType != null && cell.DataType.Value == CellValues.SharedString)
            {
                return stringTablePart.SharedStringTable.ChildElements[Int32.Parse(value)].InnerText;
            }
            else
            {
                return value;
            }
        }

【讨论】:

    【解决方案3】:

    试试这个代码,我做了一点修改,它对我有用。

    public static DataTable Fill_dataTable(string filePath)
    {
        DataTable dt = new DataTable();
    
        using (SpreadsheetDocument doc = SpreadsheetDocument.Open(filePath, false))
        {
            Sheet sheet = doc.WorkbookPart.Workbook.Sheets.GetFirstChild<Sheet>();
            Worksheet worksheet = doc.WorkbookPart.GetPartById(sheet.Id.Value) as WorksheetPart.Worksheet;
            IEnumerable<Row> rows = worksheet.GetFirstChild<SheetData>().Descendants<Row>();
            DataTable dt = new DataTable();
            List<string> columnRef = new List<string>();
            foreach (Row row in rows)
            {
                if (row.RowIndex != null)
                {
                    if (row.RowIndex.Value == 1)
                    {
                        foreach (Cell cell in row.Descendants<Cell>())
                        {
                            dt.Columns.Add(GetValue(doc, cell));
                                columnRef.Add(cell.CellReference.ToString().Substring(0, cell.CellReference.ToString().Length - 1));
                         }
                    }
                    else
                    {
                        dt.Rows.Add();
                        int i = 0;
                        foreach (Cell cell in row.Descendants<Cell>())
                        {
                            while (columnRef(i) + dt.Rows.Count + 1 != cell.CellReference)
                            {
                                dt.Rows(dt.Rows.Count - 1)(i) = "";
                                i += 1;
                             }
    
                             dt.Rows(dt.Rows.Count - 1)(i) = GetValue(doc, cell);
                             i += 1;
                        }
                    }
                }
            }
        }
    
        return dt;
    }
    

    【讨论】:

      猜你喜欢
      • 2020-05-17
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-03-20
      • 2011-04-19
      • 1970-01-01
      • 2021-05-10
      相关资源
      最近更新 更多