【问题标题】:Different output from Libtorch C++ and pytorchLibtorch C++ 和 pytorch 的不同输出
【发布时间】:2020-12-09 15:17:00
【问题描述】:

我在 pytorch 和 libtorch 中使用相同的跟踪模型,但我得到不同的输出。

Python 代码:

import cv2
import numpy as np 
import torch
import torchvision
from torchvision import transforms as trans


# device for pytorch
device = torch.device('cuda:0')

torch.set_default_tensor_type('torch.cuda.FloatTensor')

model = torch.jit.load("traced_facelearner_model_new.pt")
model.eval()

# read the example image used for tracing
image=cv2.imread("videos/example.jpg")

test_transform = trans.Compose([
        trans.ToTensor(),
        trans.Normalize([0.5, 0.5, 0.5], [0.5, 0.5, 0.5])
    ])       

resized_image = cv2.resize(image, (112, 112))

tens = test_transform(resized_image).to(device).unsqueeze(0)
output = model(tens)
print(output)

C++ 代码:

#include <iostream>
#include <algorithm> 
#include <opencv2/opencv.hpp>
#include <torch/script.h>


int main()
{
    try
    {
        torch::jit::script::Module model = torch::jit::load("traced_facelearner_model_new.pt");
        model.to(torch::kCUDA);
        model.eval();

        cv::Mat visibleFrame = cv::imread("example.jpg");

        cv::resize(visibleFrame, visibleFrame, cv::Size(112, 112));
        at::Tensor tensor_image = torch::from_blob(visibleFrame.data, { 1, visibleFrame.rows, 
                                                    visibleFrame.cols, 3 }, at::kByte);
        tensor_image = tensor_image.permute({ 0, 3, 1, 2 });
        tensor_image = tensor_image.to(at::kFloat);

        tensor_image[0][0] = tensor_image[0][0].sub(0.5).div(0.5);
        tensor_image[0][1] = tensor_image[0][1].sub(0.5).div(0.5);
        tensor_image[0][2] = tensor_image[0][2].sub(0.5).div(0.5);

        tensor_image = tensor_image.to(torch::kCUDA);
        std::vector<torch::jit::IValue> input;
        input.emplace_back(tensor_image);
        // Execute the model and turn its output into a tensor.
        auto output = model.forward(input).toTensor();
        output = output.to(torch::kCPU);
        std::cout << "Embds: " << output << std::endl;

        std::cout << "Done!\n";
    }
    catch (std::exception e)
    {
        std::cout << "exception" << e.what() << std::endl;
    }
}

模型给出(1x512)大小的输出张量如下图。

Python 输出

tensor([[-1.6270e+00, -7.8417e-02, -3.4403e-01, -1.5171e+00, -1.3259e+00,

-1.1877e+00, -2.0234e-01, -1.0677e+00, 8.8365e-01, 7.2514e-01,

2.3642e+00, -1.4473e+00, -1.6696e+00, -1.2191e+00, 6.7770e-01,

...

-7.1650e-01, 1.7661e-01]], device=‘cuda:0’,
grad_fn=)

C++ 输出

Embds: Columns 1 to 8 -84.6285 -14.7203 17.7419 47.0915 31.8170 57.6813 3.6089 -38.0543


Columns 9 to 16 3.3444 -95.5730 90.3788 -10.8355 2.8831 -14.3861 0.8706 -60.7844

...

Columns 505 to 512 36.8830 -31.1061 51.6818 8.2866 1.7214 -2.9263 -37.4330 48.5854

[ CPUFloatType{1,512} ]

使用

  • Pytorch 1.6.0
  • Libtorch 1.6.0
  • 视觉工作室 2019
  • Windows 10
  • 库达 10.1

【问题讨论】:

  • 您比我们更了解您的(相当长的)代码。如果您希望我们帮助您,最好提供您对这个问题的想法。为什么你认为这段代码没有提供正确的输出?这段代码实际上应该做什么?
  • c++ 和 python 代码本质上都在做同样的事情,即加载一个 CNN 模型并为其提供输入。如上所述的输出是 (1x512) 大小的张量。问题是模型给出的这个输出张量中的值在 C++ 和 python 中是不同的。我不确定为什么会发生这种情况,即使输入图像、预处理步骤、模型在两者中都是相同的。
  • 你只需要缩放一次 "tensor_image.sub_(0.5).div_(0.5); 在创建张量后也尝试解压你的张量,(从 load_from_blob 中删除 1 并简单地使用相应的行和列)你也不需要 IValue ,只需使用model.forward({tensor_image})
  • 顺便说一下,在此之前,您需要将张量重新缩放 255。然后进行归一化

标签: c++ pytorch jit libtorch


【解决方案1】:

在最终归一化之前,您需要将输入缩放到 0-1 范围,然后继续您正在执行的归一化。转换为浮点数,然后除以 255 应该可以到达那里。这是我写的 sn-p,可能有一些语法错误,应该是可见的。
试试这个:

#include <iostream>
#include <algorithm> 
#include <opencv2/opencv.hpp>
#include <torch/script.h>


int main()
{
    try
    {
        torch::jit::script::Module model = torch::jit::load("traced_facelearner_model_new.pt");
        model.to(torch::kCUDA);
        
        cv::Mat visibleFrame = cv::imread("example.jpg");

        cv::resize(visibleFrame, visibleFrame, cv::Size(112, 112));
        at::Tensor tensor_image = torch::from_blob(visibleFrame.data, {  visibleFrame.rows, 
                                                    visibleFrame.cols, 3 }, at::kByte);
        
        tensor_image = tensor_image.to(at::kFloat).div(255).unsqueeze(0);
        tensor_image = tensor_image.permute({ 0, 3, 1, 2 });
        ensor_image.sub_(0.5).div_(0.5);

        tensor_image = tensor_image.to(torch::kCUDA);
        // Execute the model and turn its output into a tensor.
        auto output = model.forward({tensor_image}).toTensor();
        output = output.cpu();
        std::cout << "Embds: " << output << std::endl;

        std::cout << "Done!\n";
    }
    catch (std::exception e)
    {
        std::cout << "exception" << e.what() << std::endl;
    }
}

我无权访问系统来运行它,所以如果您遇到以下任何评论。

【讨论】:

  • 非常感谢!这行得通。但是你能告诉我为什么在标准化之前将张量重新缩放 255 吗?
  • @Arki99,这是 pytorchs ToTensor 的默认设置。当您在 Pytorch 转换中执行 ToTensor() 时,它只是将输入图像重新缩放到 0-1 的范围。所以为了得到同样的行为,你需要在 libtorch 中做同样的事情。
  • 知道了。非常感谢。
猜你喜欢
  • 2020-08-19
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-09-09
  • 2020-12-28
  • 1970-01-01
  • 2021-05-01
  • 1970-01-01
相关资源
最近更新 更多