【问题标题】:Performance of C++/CLI function pointers versus .NET delegatesC++/CLI 函数指针与 .NET 委托的性能
【发布时间】:2012-11-06 18:03:03
【问题描述】:

对于我的 C++/CLI 项目,我只是尝试衡量 C++/CLI 函数指针与 .NET 委托的成本。

我的期望是,C++/CLI 函数指针比 .NET 委托更快。 所以我的测试分别统计了 5 秒内 .NET 委托和本机函数指针的调用次数。

结果

现在结果让我(现在仍然)震惊:

  • .NET 委托: 910M 次执行,结果 152080413333030 在 5003 毫秒内
  • 函数指针: 在 5013 毫秒内执行 347M 次,结果为 57893422166551

这意味着,本机 C++/CLI 函数指针的使用速度几乎比在 C++/CLI 代码中使用托管委托慢 3 倍。 这怎么可能?在性能关键部分使用接口、委托或抽象类时,我应该使用托管构造吗?

测试代码

被连续调用的函数:

__int64 DoIt(int n, __int64 sum)
{
    if ((n % 3) == 0)
        return sum + n;
    else
        return sum + 1;
}

调用该方法的代码尝试使用所有参数以及返回值,因此没有任何东西被优化掉(希望如此)。这是代码(用于 .NET 委托):

__int64 executions;
__int64 result;
System::Diagnostics::Stopwatch^ w = gcnew System::Diagnostics::Stopwatch();

System::Func<int, __int64, __int64>^ managedPtr = gcnew System::Func<int, __int64, __int64>(&DoIt);
w->Restart();
executions = 0;
result = 0;
while (w->ElapsedMilliseconds < 5000)
{
    for (int i=0; i < 1000000; i++)
        result += managedPtr(i, executions);
    executions++;
}
System::Console::WriteLine(".NET delegate:       {0}M executions with result {2} in {1}ms", executions, w->ElapsedMilliseconds, result);

与 .NET 委托调用类似,使用 C++ 函数指针:

typedef __int64 (* DoItMethod)(int n, __int64 sum);

DoItMethod nativePtr = DoIt;
w->Restart();
executions = 0;
result = 0;
while (w->ElapsedMilliseconds < 5000)
{
    for (int i=0; i < 1000000; i++)
        result += nativePtr(i, executions);
    executions++;
}
System::Console::WriteLine("Function pointer:    {0}M executions with result {2} in {1}ms", executions, w->ElapsedMilliseconds, result);

其他信息

  • 使用 Visual Studio 2012 编译
  • .NET Framework 4.5 被定位
  • 发布版本(执行计数与调试版本保持成比例)
  • 调用约定是 __stdcall(当项目使用 CLR 支持编译时不允许使用 __fastcall)

所有测试完成:

  • .NET 虚拟方法:在 5004 毫秒内执行 1025M,结果为 171358304166325
  • .NET 委托:在 5003 毫秒内执行 910M,结果为 152080413333030
  • 虚拟方法:在 5006 毫秒内执行 336M,结果为 56056335999888
  • 函数指针:在 5013 毫秒内执行 347M,结果为 57893422166551
  • 函数调用:在 5001 毫秒内执行 1459M 次,结果为 244230520832847
  • 内联函数:在 5000 毫秒内执行 1385M 次,结果为 231791984166205

此处对“DoIt”的直接调用由“函数调用”表示,它似乎被编译器内联,因为与调用内联函数相比,执行计数没有(显着)差异。

对 C++ 虚方法的调用与函数指针一样“慢”。托管类(引用类)的虚拟方法与 .NET 委托一样快。

更新: 我挖得更深一点,似乎对于使用非托管函数的测试,每次调用 DoIt 函数时都会转换到本机代码。 因此,我将内部循环包装到另一个我强制编译非托管的函数中:

#pragma managed(push, off)
__int64 TestCall(__int64* executions)
{
    __int64 result = 0;
    for (int i=0; i < 1000000; i++)
            result += DoItNative(i, *executions);
    (*executions)++;
    return result;
}
#pragma managed(pop)

另外我测试了 std::function 这样的:

#pragma managed(push, off)
__int64 TestStdFunc(__int64* executions)
{
    __int64 result = 0;
    std::function<__int64(int, __int64)> func(DoItNative);
    for (int i=0; i < 1000000; i++)
        result += func(i, *executions);
    (*executions)++;
    return result;
}
#pragma managed(pop)

现在,新结果是:

  • 函数调用:在 5000 毫秒内执行 2946M,结果为 495340439997054
  • std::function: 160M 次执行,结果 26679519999840 在 5018 毫秒内

std::function 有点令人失望。

【问题讨论】:

  • 对于这么小的函数,当内联函数比函数调用慢时,这些数字有些奇怪。
  • 假设函数调用也是内联的,我认为这在容差范围内。
  • @Hans:我认为这个问题表明 C++/CLI 函数指针的性能不如原生 C++ 函数指针,而且它确实是在询问前者。

标签: .net performance delegates c++-cli mixed-mode


【解决方案1】:

您看到了“双重打击”的成本。 DoIt() 函数的核心问题是它被编译为托管代码。委托调用非常快,通过委托从托管代码转到托管代码并不复杂。函数指针很慢,但是编译器会自动生成代码,首先从托管代码切换到非托管代码,然后通过函数指针进行调用。然后最终在一个存根中,从非托管代码切换回托管代码并调用 DoIt()。

大概您真正想要衡量的是对本机代码的调用。使用 #pragma 强制将 DoIt() 生成为机器码,如下所示:

#pragma managed(push, off)
__int64 DoIt(int n, __int64 sum)
{
    if ((n % 3) == 0)
        return sum + n;
    else
        return sum + 1;
}
#pragma managed(pop)

您现在将看到函数指针比委托更快

【讨论】:

  • 太好了,我猜它会是这种性质的东西。很好的解释!
  • 这完美地回答了这个问题!现在函数指针的执行计数为1260M。有趣的是,编译器似乎绝对需要这个提示。我通过定义 DoIt 方法两次对此进行了测试,一次用作函数指针,一次用作委托,这没有帮助:两者都 - 正如你所说 - 编译为托管代码。
  • @uebe:出于好奇,你能不能也试试在这里测量 std::function 的性能,en.cppreference.com/w/cpp/utility/functional/function
猜你喜欢
  • 2012-01-30
  • 1970-01-01
  • 2010-09-28
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2011-04-11
  • 1970-01-01
相关资源
最近更新 更多