【问题标题】:Conversion from ASCII to Unicode char code (FreeType2)从 ASCII 转换为 Unicode 字符代码 (FreeType2)
【发布时间】:2012-10-06 16:40:08
【问题描述】:

我在我的一个项目中使用 FreeType2。为了呈现一个字母,我需要提供一个 Unicode 两字节字符代码。程序读取的字符代码是 ASCII 单字节格式。 128以下的字符码没有问题(字符码相同),但其他128不匹配。例如:

ASCII 中的“a”是 0x61,Unicode 中的“a”是 0x0061 - 没关系
ASCII 中的“±”是 0xB9,Unicode 中的“±”是 0x0105 - 完全不同

我试图在那里使用 WinAPI 函数,但我一定是做错了什么。这是一个示例:

unsigned char szTest1[] = "ąółź"; //ASCII format
wchar_t* wszTest2;
int size = MultiByteToWideChar(CP_UTF8, 0, (char*)szTest1, 4, NULL, 0);
printf("size = %d\n", size);
wszTest2 = new wchar_t[size];
MultiByteToWideChar(CP_UTF8, 0, (char*)szTest1, 4, wszTest2, size);
printf("HEX: %x\n", wszTest2[0]);
delete[] wszTest2;

我希望创建一个新的宽字符串,最后没有 NULL。但是,size 变量始终等于 0。知道我做错了什么吗?或者也许有更简单的方法来解决这个问题?

【问题讨论】:

  • 没有ASCII CODES 127 以上!根据定义!对不起喊了。现在我已经引起了您的注意,您需要找出您的“ascii 文本”的实际编码是什么,并实际使用该编码对其进行解码。跨度>

标签: c++ unicode ascii freetype


【解决方案1】:

“纯”ASCII 字符集限制在 0-127(7 位)范围内。设置了最高有效位的 8 位字符(即 128-255 范围内的字符)不是唯一定义的:它们的定义取决于 代码页。 因此,您的角色 ą带有 OGONEK 的拉丁小写字母 A)由 特定 代码页中的值 0xB9 表示,该值应为 Windows-1250。在其他代码页中,值0xB9不同 字符相关联(例如,在Windows 1252 code page 中,0xB9 与字符¹ 相关联,即上标数字1)。

要使用 Windows Win32 API 将字符从特定代码页转换为 Unicode UTF-16,您可以使用 MultiByteToWideChar,指定正确的代码页(不是CP_UTF8在您问题的代码中;实际上,CP_UTF8 标识 Unicode UTF-8)。您可能想尝试将1250(ANSI 中欧;中欧(Windows))指定为正确的code page identifier

如果您可以在代码中访问 ATL,则可以使用 ATL string conversion helper classes 的便利性,例如 CA2W,它将 MultiByteToWideChar() 调用和内存分配包装在 RAII 类中;例如:

#include <atlconv.h> // ATL String Conversion Helpers
// 'test' is a Unicode UTF-16 string.
// Conversion is done from code-page 1250
// (ANSI Central European; Central European (Windows))
CA2W test("ąółź", 1250);

现在您应该可以在 Unicode API 中使用 test 字符串了。

如果您无权访问 ATL 或想要 基于 C++ STL 的解决方案,您可能需要考虑以下代码:

///////////////////////////////////////////////////////////////////////////////
//
// Modern STL-based C++ wrapper to Win32's MultiByteToWideChar() C API.
//
// (based on http://code.msdn.microsoft.com/windowsdesktop/C-UTF-8-Conversion-Helpers-22c0a664)
//
///////////////////////////////////////////////////////////////////////////////

#include <exception>    // for std::exception
#include <iostream>     // for std::cout
#include <ostream>      // for std::endl
#include <stdexcept>    // for std::runtime_error
#include <string>       // for std::string and std::wstring
#include <Windows.h>    // Win32 Platform SDK

//-----------------------------------------------------------------------------
// Define an exception class for string conversion error.
//-----------------------------------------------------------------------------
class StringConversionException 
    : public std::runtime_error
{
public:
    // Creates exception with error message and error code.
    StringConversionException(const char* message, DWORD error)
        : std::runtime_error(message)
        , m_error(error)
    {}

    // Creates exception with error message and error code.
    StringConversionException(const std::string& message, DWORD error)
        : std::runtime_error(message)
        , m_error(error)
    {}

    // Windows error code.
    DWORD Error() const
    {
        return m_error;
    }

private:
    DWORD m_error;
};

//-----------------------------------------------------------------------------
// Converts an ANSI/MBCS string to Unicode UTF-16.
// Wraps MultiByteToWideChar() using modern C++ and STL.
// Throws a StringConversionException on error.
//-----------------------------------------------------------------------------
std::wstring ConvertToUTF16(const std::string & source, const UINT codePage)
{
    // Fail if an invalid input character is encountered
    static const DWORD conversionFlags = MB_ERR_INVALID_CHARS;

    // Require size for destination string
    const int utf16Length = ::MultiByteToWideChar(
        codePage,           // code page for the conversion
        conversionFlags,    // flags
        source.c_str(),     // source string
        source.length(),    // length (in chars) of source string
        NULL,               // unused - no conversion done in this step
        0                   // request size of destination buffer, in wchar_t's
        );
    if (utf16Length == 0) 
    {
        const DWORD error = ::GetLastError();
        throw StringConversionException(
            "MultiByteToWideChar() failed: Can't get length of destination UTF-16 string.",
            error);
    }

    // Allocate room for destination string
    std::wstring utf16Text;
    utf16Text.resize(utf16Length);

    // Convert to Unicode UTF-16
    if ( ! ::MultiByteToWideChar(
        codePage,           // code page for conversion
        0,                  // validation was done in previous call
        source.c_str(),     // source string
        source.length(),    // length (in chars) of source string
        &utf16Text[0],      // destination buffer
        utf16Text.length()  // size of destination buffer, in wchar_t's
        )) 
    {
        const DWORD error = ::GetLastError();
        throw StringConversionException(
            "MultiByteToWideChar() failed: Can't convert to UTF-16 string.",
            error);
    }

    return utf16Text;
}

//-----------------------------------------------------------------------------
// Test.
//-----------------------------------------------------------------------------
int main()
{
    // Error codes
    static const int exitOk = 0;
    static const int exitError = 1;

    try 
    {
        // Test input string:
        //
        // ą - LATIN SMALL LETTER A WITH OGONEK
        std::string inText("x - LATIN SMALL LETTER A WITH OGONEK");
        inText[0] = 0xB9;

        // ANSI Central European; Central European (Windows) code page
        static const UINT codePage = 1250;

        // Convert to Unicode UTF-16
        const std::wstring utf16Text = ConvertToUTF16(inText, codePage);

        // Verify conversion.
        //  ą - LATIN SMALL LETTER A WITH OGONEK
        //  --> Unicode UTF-16 0x0105
        // http://www.fileformat.info/info/unicode/char/105/index.htm
        if (utf16Text[0] != 0x0105) 
        {
            throw std::runtime_error("Wrong conversion.");
        }
        std::cout << "All right." << std::endl;
    }
    catch (const StringConversionException& e)
    {
        std::cerr << "*** ERROR:\n";
        std::cerr << e.what() << "\n";
        std::cerr << "Error code = " << e.Error();
        std::cerr << std::endl;
        return exitError;
    }
    catch (const std::exception& e)
    {
        std::cerr << "*** ERROR:\n";
        std::cerr << e.what();
        std::cerr << std::endl;
        return exitError;
    }
    return exitOk;
}

///////////////////////////////////////////////////////////////////////////////

【讨论】:

    【解决方案2】:

    MultiByteToWideCharCodePage 参数错误。 utf-8 与 ASCII 不同。您应该使用CP_ACP,它表示当前系统代码页(与 ASCII 不同 - 请参阅 Unicode, UTF, ASCII, ANSI format differences

    大小很可能为零,因为您的测试字符串不是有效的 Utf-8 字符串。

    对于几乎所有 Win32 函数,您都可以在函数无法获取详细错误代码后调用 GetLastError(),因此调用它也会为您提供更多详细信息。

    【讨论】:

    • 是的,测试字符串绝对不是有效的 UTF-8:0xB9 不是有效的前导字节。
    • 是的,确实有一个错误。它是 ERROR_NO_UNICODE_TRANSLATION。在将 CP_UTF8 更改为 CP_ACP 后,size 变量增加为 4,正如所怀疑的那样。此外,返回了正确的字符代码,太好了!还要感谢编码格式之间的差异列表。会派上用场的。干杯
    猜你喜欢
    • 2018-03-21
    • 2018-11-19
    • 2018-11-18
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-03-29
    • 2013-01-26
    相关资源
    最近更新 更多