我遇到了两个导致延迟作业静默失败的原因。第一个是人们在分叉进程中使用 libxml 时的实际段错误(这在一段时间前出现在邮件列表中)。
第二个问题与delayed_job所依赖的1.1.0版本的守护进程有问题(https://github.com/collectiveidea/delayed_job/issues#issue/81),这可以通过使用1.0.10轻松解决,这是我自己的Gemfile所拥有的它。
记录
delayed_job 有日志记录,所以如果工作人员在没有打印错误的情况下死亡,通常是因为它没有抛出异常(例如 Segfault)或外部因素正在杀死进程。
监控
我使用 bluepill 来监控我的延迟作业实例,到目前为止,这在确保作业保持运行方面非常成功。为应用程序运行 bluepill 的步骤非常简单
将 bluepill gem 添加到您的 Gemfile:
# Monitoring
gem 'i18n' # Not sure why but it complained I didn't have it
gem 'bluepill'
我创建了一个 bluepill 配置文件:
app_home = "/home/mi/production"
workers = 5
Bluepill.application("mi_delayed_job", :log_file => "#{app_home}/shared/log/bluepill.log") do |app|
(0...workers).each do |i|
app.process("delayed_job.#{i}") do |process|
process.working_dir = "#{app_home}/current"
process.start_grace_time = 10.seconds
process.stop_grace_time = 10.seconds
process.restart_grace_time = 10.seconds
process.start_command = "cd #{app_home}/current && RAILS_ENV=production ruby script/delayed_job start -i #{i}"
process.stop_command = "cd #{app_home}/current && RAILS_ENV=production ruby script/delayed_job stop -i #{i}"
process.pid_file = "#{app_home}/shared/pids/delayed_job.#{i}.pid"
process.uid = "mi"
process.gid = "mi"
end
end
end
然后在我刚刚添加的 capistrano 部署文件中:
# Bluepill related tasks
after "deploy:update", "bluepill:quit", "bluepill:start"
namespace :bluepill do
desc "Stop processes that bluepill is monitoring and quit bluepill"
task :quit, :roles => [:app] do
run "cd #{current_path} && bundle exec bluepill --no-privileged stop"
run "cd #{current_path} && bundle exec bluepill --no-privileged quit"
end
desc "Load bluepill configuration and start it"
task :start, :roles => [:app] do
run "cd #{current_path} && bundle exec bluepill --no-privileged load /home/mi/production/current/config/delayed_job.bluepill"
end
desc "Prints bluepills monitored processes statuses"
task :status, :roles => [:app] do
run "cd #{current_path} && bundle exec bluepill --no-privileged status"
end
end
希望这会有所帮助。