【问题标题】:How to determine probability of words?如何确定单词的概率?
【发布时间】:2015-10-08 23:08:37
【问题描述】:

我有两份文件。 Doc1 格式如下:

TOPIC:  0 5892.0
site 0.0371690427699
Internet 0.0261371350984
online 0.0229124236253
web 0.0218940936864
say 0.0159538357094

TOPIC:  1 12366.0
web 0.150331554262
site 0.0517548115801
say 0.0451237263464
Internet 0.0153647096879
online 0.0135856380398

...依此类推,直到主题 99 以相同的模式。

而Doc2的格式是:

0 0.566667 0 0.0333333 0 0 0 0.133333 ..........

等等...每个主题的每个值总共有 100 个值。

现在,我要找到每个单词的加权平均概率,即:

P(w) = alpha.P(w1)+ alpha.P(w2)+...... +alpha.P(wn)

where alpha = value in the nth position corresponding to the nth topic. 

也就是说对于“说”这个词,概率应该是

P(say) = 0*0.0159 + 0.5666*0.045+....... 

同样,对于每个单词,我都必须计算概率。

For  multiplication, if the word is taken from topic 0, then the 0th value from the doc2 must be considered and so on.

我只使用以下代码对单词的出现次数进行了计数,但从未计算过它们的值。所以,我很困惑。

 with open(doc2, "r") as f:
    with open(doc3, "w") as f1:

         words = " ".join(line.strip() for line in f)
         d = defaultdict(int)
         for word in words.split():  
              d[word] += 1
              for key, value in d.iteritems() :
                  f1.write(key+ ' ' + str(value) + ' ')
              print '\n'

我的输出应该是这样的:

 say = "prob of this word calculated by above formula"
 site = "
 internet = " 

等等。

我做错了什么?

【问题讨论】:

    标签: python linux probability


    【解决方案1】:

    假设您忽略 TOPIC 行,请使用 defaultdict 对值进行分组,然后在最后进行计算:

    from collections import defaultdict
    from itertools import groupby, imap
    
    d = defaultdict(list)
    with open("doc1") as f,open("doc2") as f2:
        values = map(float, f2.read().split()) 
        for line in f:
            if line.strip() and not line.startswith("TOPIC"):
                name, val = line.split()
                d[name].append(float(val))
    
    for k,v in d.items():
        print("Prob for {} is {}".format(k ,sum(i*j for i, j in zip(v,values)) ))
    

    另一种方法是边走边计算,每次点击新部分时增加计数,即带有 TOPIC 的行,以通过索引从值中获取正确值:

    from collections import defaultdict
    d = defaultdict(float)
    from itertools import  imap
    
    with open("doc1") as f,open("doc2") as f2:
        # create list of all floats from doc2
        values = imap(float, f2.read().split())
        for line in f:
            # if we have a new TOPIC increase the ind to get corresponding ndex from values
            if line.startswith("TOPIC"):
                ind = next(values)
                continue
            # ignore empty lines
            if line.strip():
                # get word and float and multiply the val by corresponding values value
                name, val = line.split()
                d[name] += float(val) * values[ind]
    
    for k,v in d.items():
        print("Prob for {} is {}".format(k ,v) )
    

    使用您的两个 doc1 内容和 doc2 内的0 0.566667 0 0.0333333 0 为两者输出以下内容:

    Prob for web is 0.085187930859
    Prob for say is 0.0255701266375
    Prob for online is 0.0076985327511
    Prob for site is 0.0293277438137
    Prob for Internet is 0.00870667394471
    

    您也可以使用 itertools groupby:

    from collections import defaultdict
    d = defaultdict(float)
    from itertools import groupby, imap
    
    with open("doc1") as f,open("doc2") as f2:
        values = imap(float, f2.read().split())
        # lambda x: not(x.strip()) will split into groups on the empty lines
        for ind, (k, v) in enumerate(groupby(f, key=lambda x: not(x.strip()))):
            if not k:
                topic = next(v) 
                #  get matching float from values
                f = next(values)
                # iterate over the group 
                for s in v:
                    name, val = s.split()
                    d[name] += (float(val) * f)
    for k,v in d.iteritems():
        print("Prob for {} is {}".format(k,v))
    

    对于 python3,所有的 itertools imaps 都应更改为 map,这也会在 python3 中返回一个迭代器。

    【讨论】:

    • so say 0.0159538357094 乘以 5892.0?我只看到您在问题中乘以 doc2 文件中的相应元素
    • 不用担心,不客气。你让我有点困惑:)
    • 没有问题,至少你自己努力解决了这个问题,这比很多人在这里做的要多;)
    • 我会在上午看看,但基本上我们只需要替换 .read().split() 分割每一行然后迭代每一行,为每一行打开一个文件并写入输出。
    • 您使用的代码是否与发布的完全相同并使用 python2?
    猜你喜欢
    • 1970-01-01
    • 2017-08-08
    • 1970-01-01
    • 1970-01-01
    • 2019-02-20
    • 2019-02-21
    • 1970-01-01
    • 2019-12-18
    • 1970-01-01
    相关资源
    最近更新 更多