【问题标题】:SET game odds simulation (MATLAB)SET 游戏赔率模拟 (MATLAB)
【发布时间】:2011-02-17 01:30:14
【问题描述】:

我最近发现了一张很棒的卡片 - SET。简而言之,共有 81 张卡片,具有四种特征:符号(椭圆形、波浪形或菱形)、颜色(红色、紫色或绿色)、数字 >(一、二或三)和阴影(实心、条纹或开放)。任务是(从选定的 12 张卡片中)找到一组 3 张卡片,其中每张卡片的四个特征中的每一个都相同,或者每张卡片都不同(不是 2+1 组合)。

我已经在 MATLAB 中对其进行了编码,以找到解决方案并估计在随机选择的卡片中出现一组的几率。

这是我估算赔率的代码:

%% initialization
K = 12; % cards to draw
NF = 4; % number of features (usually 3 or 4)
setallcards = unique(nchoosek(repmat(1:3,1,NF),NF),'rows'); % all cards: rows - cards, columns - features
setallcomb = nchoosek(1:K,3); % index of all combinations of K cards by 3

%% test
tic
NIter=1e2; % number of test iterations
setexists = 0; % test results holder
% C = progress('init'); % if you have progress function from FileExchange
for d = 1:NIter
% C = progress(C,d/NIter);    

% cards for current test
setdrawncardidx = randi(size(setallcards,1),K,1);
setdrawncards = setallcards(setdrawncardidx,:);

% find all sets in current test iteration
for setcombidx = 1:size(setallcomb,1)
    setcomb = setdrawncards(setallcomb(setcombidx,:),:);
    if all(arrayfun(@(x) numel(unique(setcomb(:,x))), 1:NF)~=2) % test one combination
        setexists = setexists + 1;
        break % to find only the first set
    end
end
end
fprintf('Set:NoSet = %g:%g = %g:1\n', setexists, NIter-setexists, setexists/(NIter-setexists))
toc

100-1000 次迭代很快,但要小心更多。在我的家用电脑上,一百万次迭代大约需要 15 个小时。无论如何,有 12 张卡片和 4 个功能,我拥有大约 13:1 的一组。这实际上是一个问题。说明书上说这个数字应该是33:1。最近由Peter Norvig 确认。他提供了 Python 代码,但我还没有测试。

那么你能找到错误吗?欢迎任何关于性能改进的 cmets。

【问题讨论】:

  • 我没有给你计算,我也不会说 matlab,但是我偶然发现了你的问题,这让我想起了我去年在 scala 中编写的 set-game,以及我想在 Freshmeat 上发表文章——但是:没时间了。现在我有时间将一些德语 vars、cmets 和消息翻译成英文,并将其放在 website for download 上;鲜肉公告仍需要几个小时。我会看看它是否适合计算页面上的集合数。

标签: matlab set simulation probability


【解决方案1】:

这是一个矢量化版本,大约一分钟可以计算出 100 万手牌。我得到了大约 28:1 的结果,所以找到“所有不同的”集合可能还有点偏差。我的猜测是,这也是您的解决方案遇到的问题。

%# initialization
K = 12; %# cards to draw
NF = 4; %# number of features (this is hard-coded to 4)
nIter = 100000; %# number of iterations

%# each card has four features. This means that a card can be represented
%# by a coordinate in 4D space. A set is a full row, column, etc in 4D
%# space. We can even parallelize the iterations, at least as long as we
%# have RAM (each hand costs 81 bytes)
%# make card space - one dimension per feature, plus one for the iterations
cardSpace = false(3,3,3,3,nIter);

%# To draw cards, we put K trues into each cardSpace. I can't think of a
%# good, fast way to draw exactly K cards that doesn't involve calling
%# unique
for i=1:nIter
    shuffle = randperm(81) + (i-1) * 81;
    cardSpace(shuffle(1:K)) = true;
end

%# to test, all we have to do is check whether there is any row, column,
%# with all 1's
isEqual = squeeze(any(any(any(all(cardSpace,1),2),3),4) | ...
    any(any(any(all(cardSpace,2),1),3),4) | ...
    any(any(any(all(cardSpace,3),2),1),4) | ...
    any(any(any(all(cardSpace,4),2),3),1));
%# to get a set of 3 cards where all symbols are different, we require that
%# no 'sub-volume' is completely empty - there may be something wrong with this
%# but since my test looked ok, I'm not going to investigate on Friday night
isDifferent = squeeze(~any(all(all(all(~cardSpace,1),2),3),4) & ...
    ~any(all(all(all(~cardSpace,1),2),4),3) & ...
    ~any(all(all(all(~cardSpace,1),3),4),2) & ...
    ~any(all(all(all(~cardSpace,4),2),3),1));

isSet = isEqual | isDifferent;

%# find the odds
fprintf('odds are %5.2f:1\n',sum(isSet)/(nIter-sum(isSet)))

【讨论】:

  • 谢谢。虽然结果看起来更好,速度也惊人,但我认为 set 的测试有问题。很难在 4D 空间中思考。 isEqual 看起来不错,这可能意味着存在 3 张卡,但只有一个功能相同。 isDifferent 我仍然没有得到。我了解子卷,但它与设定规则有何关系?如果你觉得没问题,请解释一下好吗?
  • @yuk:我应该在 3D 中完成它 - 但是,嘿,喝了几杯啤酒之后是星期五。无论如何,我不认为测试是正确的——我误解了规则。我正在测试一个特征是否相同或所有特征不同,而不是“对于每个维度,特征是相同还是不同”。我还没有想出一个逻辑(虽然我承认我没有很努力地尝试)。
【解决方案2】:

我发现了我的错误。感谢 Jonas 对 RANDPERM 的提示。

我使用 RANDI 随机抽取 K 张牌,但即使在 12 张牌中也有大约 50% 的机会重复。当我用 randperm 替换这条线时,我得到了 33.8:1 和 10000 次迭代,非常接近说明书中的数字。

setdrawncardidx = randperm(81);
setdrawncardidx = setdrawncardidx(1:K);

无论如何,看看其他解决问题的方法会很有趣。

【讨论】:

  • 可能想将此纳入原始问题。
【解决方案3】:

我确信我对这些赔率的计算有问题,因为其他几个人已经通过模拟确认它接近说明中的 33:1,但以下逻辑有什么问题?

对于 12 张随机牌,三张牌有 220 种可能的组合 (12!/(9!3!) = 220)。三张牌的每一种组合都有 1/79 的机会成为一组,因此有 78/79 的机会三张任意的牌不是一组。因此,如果您检查了所有 220 个组合,并且每个组合都不是一组的概率为 78/79,那么您没有找到一组检查所有可能组合的概率将是 78/79 的 220 次方,即 0.0606,这是大约。 17:1 赔率。

我一定是错过了什么……?

克里斯托弗

【讨论】:

  • 如果你有三张随机牌,它们有 1/79 的机会构成一组。这是因为如果你有两张随机牌,剩下的 79 张牌中,正好有一张牌会组成一组。因此,如果你在 79 次中选择了 78 个,那么你将不会有一套。
  • 抱歉很久没有回复。我想了一会儿,我认为你的逻辑是可以的,如果你从整副牌中随机选择 3 张牌 220 次。但是由于您先随机选择 12 张牌,然后仅将这些牌组合起来,因此这是行不通的。无论如何 +1 有趣的一点。
  • 虽然我无法确定,但您的担忧似乎是有道理的。所以我编写了一个 PHP 程序,它尝试多次迭代交易并检查集合。使用这种方法处理 1,000,000 笔交易,脚本计算出 96.8% 的交易有一组或多组。这与所说的交易不包含集合的 1/30 机会一致。您可以在此处运行代码(它只运行 10,000 次迭代,因为 Web 服务器在 30 秒后超时):susdesign.com/temp/set.php 这是代码的可读版本:susdesign.com/temp/set-code.html 有什么想法吗?克里斯托弗
【解决方案4】:

在查看您的代码之前,我解决了编写自己的实现的问题。我的第一次尝试与您已经拥有的非常相似:)

%# some parameters
NUM_ITER = 100000;  %# number of simulations to run
DRAW_SZ = 12;       %# number of cards we are dealing
SET_SZ = 3;         %# number of cards in a set
FEAT_NUM = 4;       %# number of features (symbol,color,number,shading)
FEAT_SZ = 3;        %# number of values per feature (eg: red/purple/green, ...)

%# cards features
features = {
    'oval' 'squiggle' 'diamond' ;    %# symbol
    'red' 'purple' 'green' ;         %# color
    'one' 'two' 'three' ;            %# number
    'solid' 'striped' 'open'         %# shading
};
fIdx = arrayfun(@(k) grp2idx(features(k,:)), 1:FEAT_NUM, 'UniformOutput',0);

%# list of all cards. Each card: [symbol,color,number,shading]
[W X Y Z] = ndgrid(fIdx{:});
cards = [W(:) X(:) Y(:) Z(:)];

%# all possible sets: choose 3 from 12
setsInd = nchoosek(1:DRAW_SZ,SET_SZ);

%# count number of valid sets in random draws of 12 cards
counterValidSet = 0;
for i=1:NUM_ITER
    %# pick 12 cards
    ord = randperm( size(cards,1) );
    cardsDrawn = cards(ord(1:DRAW_SZ),:);

    %# check for valid sets: features are all the same or all different
    for s=1:size(setsInd,1)
        %# set of 3 cards
        set = cardsDrawn(setsInd(s,:),:);

        %# check if set is valid
        count = arrayfun(@(k) numel(unique(set(:,k))), 1:FEAT_NUM);
        isValid = (count==1|count==3);

        %# increment counter
        if isValid
            counterValidSet = counterValidSet + 1;
            break           %# break early if found valid set among candidates
        end
    end
end

%# ratio of found-to-notfound
fprintf('Size=%d, Set=%d, NoSet=%d, Set:NoSet=%g\n', ...
    DRAW_SZ, counterValidSet, (NUM_ITER-counterValidSet), ...
    counterValidSet/(NUM_ITER-counterValidSet))

在使用 Profiler 发现热点之后,可以通过尽可能早地退出循环来进行一些改进。主要瓶颈是对 UNIQUE 函数的调用。上面我们检查有效集合的那两行可以重写为:

%# check if set is valid
isValid = true;
for k=1:FEAT_NUM
    count = numel(unique(set(:,k)));
    if count~=1 && count~=3
        isValid = false;
        break   %# break early if one of the features doesnt meet conditions
    end
end

不幸的是,对于较大的模拟,模拟仍然很慢。因此,我的下一个解决方案是矢量化版本,对于每次迭代,我们从 12 张抽出的牌中构建所有可能的 3 张牌组的单个矩阵。对于所有这些候选集,我们使用逻辑向量来指示存在什么特征,从而避免调用 UNIQUE/NUMEL(我们希望集合的每张卡片上的特征都相同或不同)。

我承认现在的代码可读性降低并且难以理解(因此我发布了两个版本以进行比较)。原因是我试图尽可能地优化代码,以便每个迭代循环都是完全矢量化的。这是最终代码:

%# some parameters
NUM_ITER = 100000;  %# number of simulations to run
DRAW_SZ = 12;       %# number of cards we are dealing
SET_SZ = 3;         %# number of cards in a set
FEAT_NUM = 4;       %# number of features (symbol,color,number,shading)
FEAT_SZ = 3;        %# number of values per feature (eg: red/purple/green, ...)

%# cards features
features = {
    'oval' 'squiggle' 'diamond' ;    %# symbol
    'red' 'purple' 'green' ;         %# color
    'one' 'two' 'three' ;            %# number
    'solid' 'striped' 'open'         %# shading
};
fIdx = arrayfun(@(k) grp2idx(features(k,:)), 1:FEAT_NUM, 'UniformOutput',0);

%# list of all cards. Each card: [symbol,color,number,shading]
[W X Y Z] = ndgrid(fIdx{:});
cards = [W(:) X(:) Y(:) Z(:)];

%# all possible sets: choose 3 from 12
setsInd = nchoosek(1:DRAW_SZ,SET_SZ);

%# optimizations: some calculations taken out of the loop
ss = setsInd(:);
set_sz2 = numel(ss)*FEAT_NUM/SET_SZ;
col = repmat(1:set_sz2,SET_SZ,1);
col = FEAT_SZ.*(col(:)-1);
M = false(FEAT_SZ,set_sz2);

%# progress indication
%#hWait = waitbar(0./NUM_ITER, 'Simulation...');

%# count number of valid sets in random draws of 12 cards
counterValidSet = 0;
for i=1:NUM_ITER
    %# update progress
    %#waitbar(i./NUM_ITER, hWait);

    %# pick 12 cards
    ord = randperm( size(cards,1) );
    cardsDrawn = cards(ord(1:DRAW_SZ),:);

    %# put all possible sets of 3 cards next to each other
    set = reshape(cardsDrawn(ss,:)',[],SET_SZ)';
    set = set(:);

    %# check for valid sets: features are all the same or all different
    M(:) = false;            %# if using PARFOR, it will complain about this
    M(set+col) = true;
    isValid = all(reshape(sum(M)~=2,FEAT_NUM,[]));

    %# increment counter if there is at least one valid set in all candidates
    if any(isValid)
        counterValidSet = counterValidSet + 1;
    end
end

%# ratio of found-to-notfound
fprintf('Size=%d, Set=%d, NoSet=%d, Set:NoSet=%g\n', ...
    DRAW_SZ, counterValidSet, (NUM_ITER-counterValidSet), ...
    counterValidSet/(NUM_ITER-counterValidSet))

%# close progress bar
%#close(hWait)

如果您有并行处理工具箱,您可以轻松地将普通 FOR 循环替换为并行 PARFOR(您可能希望再次将矩阵 M 的初始化移动到循环内:将 M(:) = false; 替换为 @987654327 @)

以下是 50000 次模拟的一些示例输出(PARFOR 与 2 个本地实例池一起使用):

» tic, SET_game2, toc
Size=12, Set=48376, NoSet=1624, Set:NoSet=29.7882
Elapsed time is 5.653933 seconds.

» tic, SET_game2, toc
Size=15, Set=49981, NoSet=19, Set:NoSet=2630.58
Elapsed time is 9.414917 seconds.

一百万次迭代(PARFOR 为 12,no-PARFOR 为 15):

» tic, SET_game2, toc
Size=12, Set=967516, NoSet=32484, Set:NoSet=29.7844
Elapsed time is 110.719903 seconds.

» tic, SET_game2, toc
Size=15, Set=999630, NoSet=370, Set:NoSet=2701.7
Elapsed time is 372.110412 seconds.

优势比与Peter Norvig报告的结果一致。

【讨论】:

    猜你喜欢
    • 2011-01-19
    • 2019-05-09
    • 2010-11-28
    • 2018-01-18
    • 1970-01-01
    • 1970-01-01
    • 2023-04-11
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多