【问题标题】:Backward pass in Caffe Python Layer is not called/working?Caffe Python层中的向后传递未被调用/工作?
【发布时间】:2017-03-25 05:15:48
【问题描述】:

我没有成功尝试使用 Caffe 在 Python 中实现一个简单的损失层。作为参考,我发现了几个用 Python 实现的层,包括 hereherehere

从 Caffe 文档/示例提供的 EuclideanLossLayer 开始,我无法让它工作并开始调试。即使使用这个简单的TestLayer

def setup(self, bottom, top):
    """
    Checks the correct number of bottom inputs.
    
    :param bottom: bottom inputs
    :type bottom: [numpy.ndarray]
    :param top: top outputs
    :type top: [numpy.ndarray]
    """
    
    print 'setup'

def reshape(self, bottom, top):
    """
    Make sure all involved blobs have the right dimension.
    
    :param bottom: bottom inputs
    :type bottom: caffe._caffe.RawBlobVec
    :param top: top outputs
    :type top: caffe._caffe.RawBlobVec
    """
    
    print 'reshape'
    top[0].reshape(bottom[0].data.shape[0], bottom[0].data.shape[1], bottom[0].data.shape[2], bottom[0].data.shape[3])
    
def forward(self, bottom, top):
    """
    Forward propagation.
    
    :param bottom: bottom inputs
    :type bottom: caffe._caffe.RawBlobVec
    :param top: top outputs
    :type top: caffe._caffe.RawBlobVec
    """
    
    print 'forward'
    top[0].data[...] = bottom[0].data

def backward(self, top, propagate_down, bottom):
    """
    Backward pass.
    
    :param bottom: bottom inputs
    :type bottom: caffe._caffe.RawBlobVec
    :param propagate_down:
    :type propagate_down:
    :param top: top outputs
    :type top: caffe._caffe.RawBlobVec
    """
    
    print 'backward'
    bottom[0].diff[...] = top[0].diff[...]

我无法让 Python 层正常工作。学习任务相当简单,因为我只是试图预测一个实数值是正数还是负数。对应的数据生成如下并写入LMDB:

N = 10000
N_train = int(0.8*N)
    
images = []
labels = []
    
for n in range(N):            
    image = (numpy.random.rand(1, 1, 1)*2 - 1).astype(numpy.float)
    label = int(numpy.sign(image))
        
    images.append(image)
    labels.append(label)

将数据写入 LMDB 应该是正确的,因为使用 Caffe 提供的 MNIST 数据集进行的测试显示没有问题。网络定义如下:

 net.data, net.labels = caffe.layers.Data(batch_size = batch_size, backend = caffe.params.Data.LMDB, 
                                                source = lmdb_path, ntop = 2)
 net.fc1 = caffe.layers.Python(net.data, python_param = dict(module = 'tools.layers', layer = 'TestLayer'))
 net.score = caffe.layers.TanH(net.fc1)
 net.loss = caffe.layers.EuclideanLoss(net.score, net.labels)

使用手动完成求解:

for iteration in range(iterations):
    solver.step(step)

对应的prototxt文件如下:

solver.prototxt:

weight_decay: 0.0005
test_net: "tests/test.prototxt"
snapshot_prefix: "tests/snapshot_"
max_iter: 1000
stepsize: 1000
base_lr: 0.01
snapshot: 0
gamma: 0.01
solver_mode: CPU
train_net: "tests/train.prototxt"
test_iter: 0
test_initialization: false
lr_policy: "step"
momentum: 0.9
display: 100
test_interval: 100000

train.prototxt:

layer {
  name: "data"
  type: "Data"
  top: "data"
  top: "labels"
  data_param {
    source: "tests/train_lmdb"
    batch_size: 64
    backend: LMDB
  }
}
layer {
  name: "fc1"
  type: "Python"
  bottom: "data"
  top: "fc1"
  python_param {
    module: "tools.layers"
    layer: "TestLayer"
  }
}
layer {
  name: "score"
  type: "TanH"
  bottom: "fc1"
  top: "score"
}
layer {
  name: "loss"
  type: "EuclideanLoss"
  bottom: "score"
  bottom: "labels"
  top: "loss"
}

test.prototxt:

layer {
  name: "data"
  type: "Data"
  top: "data"
  top: "labels"
  data_param {
    source: "tests/test_lmdb"
    batch_size: 64
    backend: LMDB
  }
}
layer {
  name: "fc1"
  type: "Python"
  bottom: "data"
  top: "fc1"
  python_param {
    module: "tools.layers"
    layer: "TestLayer"
  }
}
layer {
  name: "score"
  type: "TanH"
  bottom: "fc1"
  top: "score"
}
layer {
  name: "loss"
  type: "EuclideanLoss"
  bottom: "score"
  bottom: "labels"
  top: "loss"
}

我尝试追踪它,在TestLayerbackwardfoward 方法中添加调试消息,在求解过程中只调用forward 方法(注意没有执行任何测试,调用只能是相关的解决方案)。同样,我在python_layer.hpp 中添加了调试消息:

virtual void Forward_cpu(const vector<Blob<Dtype>*>& bottom,
    const vector<Blob<Dtype>*>& top) {
  LOG(INFO) << "cpp forward";
  self_.attr("forward")(bottom, top);
}
virtual void Backward_cpu(const vector<Blob<Dtype>*>& top,
    const vector<bool>& propagate_down, const vector<Blob<Dtype>*>& bottom) {
  LOG(INFO) << "cpp backward";
  self_.attr("backward")(top, propagate_down, bottom);
}

同样,只执行前向传球。当我删除TestLayer 中的backward 方法时,求解仍然有效。删除forward 方法时,由于forward 未实现而引发错误。我希望backward 也是如此,因此似乎根本没有执行向后传递。切换回常规层并添加调试消息,一切正常。

我感觉我错过了一些简单或基本的东西,但我已经好几天没能解决问题了。因此,感谢任何帮助或提示。

谢谢!

【问题讨论】:

    标签: neural-network deep-learning caffe pycaffe


    【解决方案1】:

    这是预期的行为,因为您的 python 层“下方”没有任何实际需要梯度来计算权重更新的层。 Caffe 注意到了这一点并跳过了这些层的反向计算,因为这会浪费时间。

    如果在网络初始化时需要在日志中进行反向计算,Caffe 会打印所有层。 在您的情况下,您应该看到如下内容:

    fc1 does not need backward computation.
    

    如果您在“Python”层下方放置“InnerProduct”或“Convolution”层(例如Data-&gt;InnerProduct-&gt;Python-&gt;Loss),则需要反向计算并调用您的反向方法。

    【讨论】:

      【解决方案2】:

      除了Erik B.的回答,还可以通过指定强制caffe回溯

      force_backward: true
      

      在您的网络 prototxt 中。
      有关更多信息,请参阅caffe.proto 中的 cmets。

      【讨论】:

        【解决方案3】:

        即使我按照 David Stutz 的建议设置了 force_backward: true,我的也无法正常工作。我发现herehere 忘记在目标类的索引处将最后一层的差异设置为1。

        正如 Mohit Jain 在他的 caffe-users 回答中所描述的,如果您正在使用虎斑猫进行 ImageNet 分类,那么在进行前向传递之后,您必须执行以下操作:

        net.blobs['prob'].diff[0][281] = 1   # 281 is tabby cat. diff shape: (1, 1000)
        

        请注意,您必须将 'prob' 相应地更改为最后一层的名称,通常是 softmax 和 'prob'

        这是一个基于我的例子:


        deploy.prototxt(它是基于VGG16松散的,只是为了显示文件的结构,但我没有测试它):

        name: "smaller_vgg"
        input: "data"
        force_backward: true
        input_dim: 1
        input_dim: 3
        input_dim: 224
        input_dim: 224
        layer {
          name: "conv1_1"
          type: "Convolution"
          bottom: "data"
          top: "conv1_1"
          convolution_param {
            num_output: 64
            pad: 1
            kernel_size: 3
          }
        }
        layer {
          name: "relu1_1"
          type: "ReLU"
          bottom: "conv1_1"
          top: "conv1_1"
        }
        layer {
          name: "pool1"
          type: "Pooling"
          bottom: "conv1_1"
          top: "pool1"
          pooling_param {
            pool: MAX
            kernel_size: 2
            stride: 2
          }
        }
        layer {
          name: "fc1"
          type: "InnerProduct"
          bottom: "pool1"
          top: "fc1"
          inner_product_param {
            num_output: 4096
          }
        }
        layer {
          name: "relu1"
          type: "ReLU"
          bottom: "fc1"
          top: "fc1"
        }
        layer {
          name: "drop1"
          type: "Dropout"
          bottom: "fc1"
          top: "fc1"
          dropout_param {
            dropout_ratio: 0.5
          }
        }
        layer {
          name: "fc2"
          type: "InnerProduct"
          bottom: "fc1"
          top: "fc2"
          inner_product_param {
            num_output: 1000
          }
        }
        layer {
          name: "prob"
          type: "Softmax"
          bottom: "fc2"
          top: "prob"
        }
        

        main.py:

        import caffe
        
        prototxt = 'deploy.prototxt'
        model_file = 'smaller_vgg.caffemodel'
        net = caffe.Net(model_file, prototxt, caffe.TRAIN)  # not sure if TEST works as well
        
        image = cv2.imread('tabbycat.jpg', cv2.IMREAD_UNCHANGED)
        
        net.blobs['data'].data[...] = image[np.newaxis, np.newaxis, :]
        net.blobs['prob'].diff[0, 298] = 1
        net.forward()
        backout = net.backward()
        
        # access grad from backout['data'] or net.blobs['data'].diff
        

        【讨论】:

        • 按照你的代码,为什么 net.blobs['prob'].diff[0, 298] 在 net.backward() 之后不再是 1。 net.backward() 会改变你的预设值吗?
        • @Stone 我不确定。我将此代码仅用于引导反向传播和 gradcam 的单次传递。可能是 Caffe 在每次迭代后重置diff(如果这是真的,那么我猜所有的渐变也必须被重置)。每次backward() 调用后设置net.blobs['prob'].diff[0, 298] = 1 是否修复它?
        • 在每次backward()调用之后设置net.blobs['prob'].diff[0, 298] = 1,本质上保证它的值仍然是1。我担心的是,如果Caffe在每次迭代后重置diff(如你所说),那么在net.backward() 之后就无法从net.blobs[layer_name].diff 访问毕业生。此外,如果在net.backward() 之后访问net.blobs[layer_name].diff 是正确的方法,那么最顶层probnet.blobs['prob'].diff)的梯度应该保持原样(例如net.blobs['prob'].diff[0, 298] = 1),因为梯度计算从@开始987654341@层。
        • @Stone 你是对的,访问net.blobs[layer].diffbackward() 之后。然后,我不明白为什么它会重置prob diff。如果你发现了什么,请告诉我。
        • 也许这只是我的问题,如果你没有看到 Caffe 在你的 backward() 之后重置 diff,那么一定是我错过了一些配置。谢谢!
        猜你喜欢
        • 1970-01-01
        • 2017-06-20
        • 1970-01-01
        • 2017-05-11
        • 1970-01-01
        • 2018-09-30
        • 1970-01-01
        • 1970-01-01
        • 2012-10-29
        相关资源
        最近更新 更多