【问题标题】:Google Datastore - Blob or TextGoogle 数据存储区 - Blob 或文本
【发布时间】:2010-06-04 20:46:05
【问题描述】:

在 Google 数据存储区中保留大字符串的两种可能方法是 TextBlob 数据类型。

从存储消耗的角度来看,推荐 2 个中的哪一个?从 protobuf 序列化和反序列化的角度来看,同样的问题。

【问题讨论】:

    标签: google-app-engine google-cloud-datastore


    【解决方案1】:

    两者之间没有显着的性能差异 - 只需使用最适合您的数据的一个即可。 BlobProperty 应该用于存储二进制数据(例如,str 对象),而TextProperty 应该用于存储任何文本数据(例如,unicodestr 对象)。请注意,如果您将str 存储在TextProperty 中,则它只能包含ASCII 字节(小于十六进制80 或十进制128)(与BlobProperty 不同)。

    这两个属性都派生自UnindexedProperty,您可以在source 中看到。

    这是一个示例应用程序,它证明这些 ASCII 或 UTF-8 字符串的存储开销没有区别:

    import struct
    
    from google.appengine.ext import db, webapp
    from google.appengine.ext.webapp.util import run_wsgi_app
    
    class TestB(db.Model):
        v = db.BlobProperty(required=False)
    
    class TestT(db.Model):
        v = db.TextProperty(required=False)
    
    class MainPage(webapp.RequestHandler):
        def get(self):
            self.response.headers['Content-Type'] = 'text/plain'
    
            # try simple ASCII data and a bytestring with non-ASCII bytes
            ascii_str = ''.join([struct.pack('>B', i) for i in xrange(128)])
            arbitrary_str = ''.join([struct.pack('>2B', 0xC2, 0x80+i) for i in xrange(64)])
            u = unicode(arbitrary_str, 'utf-8')
    
            t = [TestT(v=ascii_str), TestT(v=ascii_str*1000), TestT(v=u*1000)]
            b = [TestB(v=ascii_str), TestB(v=ascii_str*1000), TestB(v=arbitrary_str*1000)]
    
            # demonstrate error cases
            try:
                err = TestT(v=arbitrary_str)
                assert False, "should have caused an error: can't store non-ascii bytes in a Text"
            except UnicodeDecodeError:
                pass
            try:
                err = TestB(v=u)
                assert False, "should have caused an error: can't store unicode in a Blob"
            except db.BadValueError:
                pass
    
            # determine the serialized size of each model (note: no keys assigned)
            fEncodedSz = lambda o : len(db.model_to_protobuf(o).Encode())
            sz_t = tuple([fEncodedSz(x) for x in t])
            sz_b = tuple([fEncodedSz(x) for x in b])
    
            # output the results
            self.response.out.write("text:   1=>%dB  2=>%dB  3=>%dB\n" % sz_t)
            self.response.out.write("blob:   1=>%dB  2=>%dB  3=>%dB\n" % sz_b)
    
    application = webapp.WSGIApplication([('/', MainPage)])
    def main(): run_wsgi_app(application)
    if __name__ == '__main__': main()
    

    这是输出:

    text:   1=>172B  2=>128047B  3=>128047B
    blob:   1=>172B  2=>128047B  3=>128047B
    

    【讨论】:

    • 我不知道 Text 属性只能包含 ASCII 字节。这种认识回答了我的问题。谢谢。
    • 这不是真的 - 文本属性存储 unicode。但是,如果您将字节('raw')字符串(类型'str')分配给文本属性,它将尝试转换为使用系统默认编码的 unicode,即 ASCII。如果您不想这样做,则需要显式解码字符串。
    • 谢谢尼克。我想说TextProperty 不能存储包含非ASCII 字节的str 对象,但是(正如你所指出的)我的评论没有说清楚,所以我删除了它。
    猜你喜欢
    • 1970-01-01
    • 2018-01-11
    • 2014-10-22
    • 2012-12-01
    • 2012-05-07
    • 1970-01-01
    • 2018-10-30
    • 1970-01-01
    相关资源
    最近更新 更多