【问题标题】:Using LINQ GroupBy chains to produce a multi-level hierarchy with aggregations使用 LINQ GroupBy 链生成具有聚合的多级层次结构
【发布时间】:2019-10-05 21:24:43
【问题描述】:

我碰巧做到了这一点,并想询问有关 O(n)、订单稳定性和所涉及的枚举器实例的底层保证。

我了解每个聚合的枚举数拐点数(例如计数和持续时间)取决于子级到父级的实际分布,但至少每个记录仅在每个聚合中枚举一次,这不是真的?

在这个例子中,我们有 3 个聚合和 3 个层次级别要聚合,因此 O(9n) ~ O(n)。

LINQ GroupBy 问题:

  1. GroupBy 是否记录为既具有线性复杂性又具有稳定的顺序,还是 Microsoft .NET 的 impl 是如何做到的?我没有特别怀疑它可能是非线性的,只是出于好奇而询问。
  2. 似乎实例化的枚举数应该始终与层次结构中看到的父节点数成线性关系,对吗?
  3. 在同一聚合统计信息和级别期间,没有多次枚举叶或非叶,对吗?因为每个节点只是某个“结果”中一个父节点的子节点。
    using System;
    using System.Collections.Generic;
    using System.Linq;
    using Microsoft.VisualStudio.TestTools.UnitTesting;

    namespace NestedGroupBy
    {
        [TestClass]
        public class UnitTest1
        {
            struct Record
            {
                // cat->subcat->event is the composite 3-tier key
                private readonly string category;
                private readonly string subcategory;
                private readonly string ev;
                private readonly double duration;

                public Record(string category, string subcategory, string ev, double duration)
                {
                    this.category = category;
                    this.subcategory = subcategory;
                    this.ev = ev;
                    this.duration = duration;
                }

                public string Category { get { return category; } }
                public string Subcategory { get { return subcategory; } }
                public string Event { get { return ev; } }
                public double Duration { get { return duration; } }
            }

            static readonly IList<Record> Rec1 = new System.Collections.ObjectModel.ReadOnlyCollection<Record>
            (new[] {
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "Security", "ReadSocket", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("File", "ReadMetadata", "ReadDirectory", 0.0145),
                 new Record ("File", "ReadMetadata", "ReadDirectory", 0.0145),
                 new Record ("File", "ReadMetadata", "ReadDirectory", 0.0145),
                 new Record ("File", "ReadMetadata", "ReadSize", 0.0145),
                 new Record ("File", "ReadMetadata", "ReadSize", 0.0145),
                 new Record ("File", "ReadMetadata", "ReadSize", 0.0145),
                 new Record ("Registry", "ReadKey", "ReadKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "ReadKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "ReadKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "ReadKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "CacheKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "CacheKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "CheckKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "CheckKeyAcl", 0.0145),
                 new Record ("Registry", "ReadKey", "CheckKeyAcl", 0.0145),
                 new Record ("Registry", "WriteKey", "CheckPermissions", 0.0145),
                 new Record ("Registry", "WriteKey", "CheckOwner", 0.0145),
                 new Record ("Registry", "WriteKey", "InheritPermissions", 0.0145),
                 new Record ("Registry", "WriteKey", "ValidateKey", 0.0145),
                 new Record ("Registry", "WriteKey", "RecacheKey", 0.0145),
                 new Record ("File", "WriteData", "FlushData", 0.0145),
                 new Record ("File", "WriteData", "WriteBuffer", 0.0145),
                 new Record ("File", "WriteData", "WritePermissions", 0.0145),
                 new Record ("File", "ReadData", "CheckDataBuffer", 0.0145),
                 new Record ("File", "ReadData", "ReadBuffer", 0.0145),
                 new Record ("File", "ReadData", "ReadPermissions", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "Security", "ReadSocket", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
                 new Record ("Network", "Security", "ReadSocket", 0.0145),
                 new Record ("Network", "SecurityCheck", "ReadAcl", 0.0145),
            });

            [TestMethod]
            public void TestMethod1()
            {
                // Perform one big sort to order all child rungs properly
                var s = Rec1.OrderBy(
                    r => r,
                    Comparer<Record>.Create(
                        (x, y) =>
                        {
                            int c = x.Category.CompareTo(y.Category);
                            if (c != 0) return c;
                            c = x.Subcategory.CompareTo(y.Subcategory);
                            if (c != 0) return c;
                            return x.Event.CompareTo(y.Event);
                        }));

                // This query enumerates bottom-up (in the key hierarchy-sense), 
                // then proceedes to each higher summary (parent) level and retains
                // the "result" collection of its children determined by the preceding GroupBy.
                //
                // This is so each level can later step down into its own children for looping. 
                // And the leaf durations, immediate child counts and leaf event counts are already calculated as well.
                //
                // I think this is O(n), since each record does not get repeatedly scanned for different levels of the same accumulation stat.
                // But under-the-hood there may be much grainy processing like enumerator instantiation, depending on child count density.
                var q = s
                    .GroupBy(
                        r => new { Category = r.Category, Subcategory = r.Subcategory, Event = r.Event },
                        (key, result) => {
                            int c = result.Count();
                            return new
                            {
                                LowKey = key,
                                // at this lowest summary level only, 
                                // the hierarchical (immediate child) count is the same as the event (leaf) count
                                LowChildCount = c,
                                LowEventCount = c,
                                LowDuration = result.Sum(x => x.Duration),
                                LowChildren = result
                            };
                        })
                        .GroupBy(
                            r => new { Category = r.LowKey.Category, Subcategory = r.LowKey.Subcategory },
                                (key, result) => new {
                                    MidKey = key,
                                    MidChildCount = result.Count(),
                                    MidEventCount = result.Sum(x => x.LowEventCount),
                                    MidDuration = result.Sum(x => x.LowDuration),
                                    MidChildren = result
                                })
                                .GroupBy(
                                    r => new { Category = r.MidKey.Category },
                                    (key, result) => new {
                                        HighKey = key,
                                        HighChildCount = result.Count(),
                                        HighEventCount = result.Sum(x => x.MidEventCount),
                                        HighDuration = result.Sum(x => x.MidDuration),
                                        HighChildren = result
                                    });


                foreach (var high in q)
                {
                    Console.WriteLine($"{high.HighKey.Category} child#={high.HighChildCount} event#={high.HighEventCount} duration={high.HighDuration}");
                    foreach (var mid in high.HighChildren)
                    {
                        Console.WriteLine($"  {mid.MidKey.Subcategory} child#={mid.MidChildCount} event#={high.HighEventCount} duration={mid.MidDuration}");
                        foreach (var low in mid.MidChildren)
                        {
                            Console.WriteLine($"    {low.LowKey.Event} child#={low.LowChildCount} event#={high.HighEventCount} duration={low.LowDuration}");
                            foreach (var leaf in low.LowChildren)
                            {
                                Console.WriteLine($"      >> {leaf.Category}/{leaf.Subcategory}/{leaf.Event} duration={leaf.Duration}");
                            }
                        }
                    }
                }
            }
        }
    }

【问题讨论】:

  • 查看GroupBy的源代码,这将回答您的所有问题
  • 谢谢,这一切都说得通。

标签: c# linq


【解决方案1】:

让我们从MS源代码中回顾Enumerable.GroupBy的实现开始:

分组方式

public static IEnumerable<TResult> GroupBy<TSource, TKey, TElement, TResult>
(this IEnumerable<TSource> source, 
 Func<TSource, TKey> keySelector, 
 Func<TSource, TElement> elementSelector, 
 Func<TKey, IEnumerable<TElement>, TResult> resultSelector)
{
  return new GroupedEnumerable<TSource, TKey, TElement, TResult>(source, 
                                                                 keySelector, 
                                                                 elementSelector, 
                                                                 resultSelector, null);
}

GroupedEnumerable

internal class GroupedEnumerable<TSource, TKey, TElement, TResult> : IEnumerable<TResult>{
    IEnumerable<TSource> source;
    Func<TSource, TKey> keySelector;
    Func<TSource, TElement> elementSelector;
    IEqualityComparer<TKey> comparer;
    Func<TKey, IEnumerable<TElement>, TResult> resultSelector;

    public GroupedEnumerable(IEnumerable<TSource> source, Func<TSource, TKey> keySelector, Func<TSource, TElement> elementSelector, Func<TKey, IEnumerable<TElement>, TResult> resultSelector, IEqualityComparer<TKey> comparer){
        if (source == null) throw Error.ArgumentNull("source");
        if (keySelector == null) throw Error.ArgumentNull("keySelector");
        if (elementSelector == null) throw Error.ArgumentNull("elementSelector");
        if (resultSelector == null) throw Error.ArgumentNull("resultSelector");
        this.source = source;
        this.keySelector = keySelector;
        this.elementSelector = elementSelector;
        this.comparer = comparer;
        this.resultSelector = resultSelector;
    }

    public IEnumerator<TResult> GetEnumerator(){
        Lookup<TKey, TElement> lookup = Lookup<TKey, TElement>.Create<TSource>(source, keySelector, elementSelector, comparer);
        return lookup.ApplyResultSelector(resultSelector).GetEnumerator();
    }

    IEnumerator IEnumerable.GetEnumerator(){
        return GetEnumerator();
    }
}

重要的一点是它在内部使用LookUp数据结构,它为Key提供O(1)查找,因此它在内部枚举所有记录并为每条记录添加数据到LookUp,这是输入IEnumerable&lt;IGrouping&lt;TKey,TValue&gt;&gt;

现在您的具体问题

GroupBy 是否记录了线性复杂性和稳定排序,或者仅仅是 Microsoft .NET 的 impl 是如何做到的?我没有特别怀疑它可能是非线性的,只是出于好奇。

设计和代码建议在顶层 O(N),但肯定取决于元素选择,如果进一步执行 O(N) 操作,那将自动使复杂度 O(N^2) 或更多

似乎实例化的枚举数应该始终与层次结构中看到的父节点数成线性关系,对吗?

是的,它只是顶层的单个枚举器

在同一聚合统计和级别期间,没有多次枚举叶子或非叶子,对吗?因为每个节点只是某个“结果”中一个父节点的子节点。

通过非叶节点,我假设您指的是嵌套数据,因此这取决于您的模型设计,但如果每个元素都按预期具有嵌套数据的单独副本,而不是共享/引用,那么我不看看为什么一个元素会被多次触摸,它是特定子/子数据的枚举

更多细节

  • 在您的情况下,您正在编写GroupBy 查询,因此一个查询的结果充当另一个查询的输入,就复杂性而言,这更像O(N) + O(N) + O(N) ~ O(N),但在每个查询中您正在执行至少 1 O(N) 操作,因此您的嵌套设计的整体复杂性应为 O(N^2) 而不是 O(N)
  • 如果您有深层嵌套数据,则改为使用串联的GroupBySelectMany,将数据展平是更好的选择,因为在展平的数据上运行最终聚合不太复杂

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-06-17
    • 1970-01-01
    • 2014-07-23
    • 1970-01-01
    相关资源
    最近更新 更多