Repository navigation
Getting a buffer from a Unicode array uses invalid format #57281
Description
Activity
In Python 3.2, when you get a buffer from array.array('u'), "u" is used as buffer format. The format is supposed to be a format from the struct module, and "u" is an invalid struct format. "w" is used on wide mode.
I just upgraded the array module to use the new Unicode API (PEP-393). The array now uses a Py_UCS4 buffer. I used "I" or "L" format depending on the size of int and long C types.
It would be better to use a format for a Py_UCS4 string, but struct doesn't support such type.
For Python 2.7 and 3.2, I don't know if it is really a bug or not.
- addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on Sep 30, 2011 The automatic conversion of 'u' to 'I' or 'L' causes test_buffer
(PEP-3118 repo) to fail:# Not implemented formats. Ugly, but inevitable. This is the same as # issue python/cpython#46783: equality is also used for membership testing and must # return a result. a = array.array('u', 'xyz') v = memoryview(a) self.assertNotEqual(v, a) self.assertNotEqual(a, v)
I don't have a better idea though what to do about 'u' except
officially implementing it for struct and memoryview as well.It would be better to use a format for a Py_UCS4 string, but struct doesn't support such type.
PEP-3118 suggests for the extended struct syntax:
'c' -> ucs-1 (latin-1) encoding
'u' -> ucs-2
'w' -> ucs-4The automatic conversion of 'u' to 'I' or 'L' causes test_buffer
(PEP-3118 repo) to fail:Not implemented formats. Ugly, but inevitable. This is the same as
issue bpo-2531: equality is also used for membership testing and must
return a result.
a = array.array('u', 'xyz')
v = memoryview(a)
self.assertNotEqual(v, a)
self.assertNotEqual(a, v)I don't understand: a buffer format is a format for the struct module,
or for the array module?STINNER Victor <report@bugs.python.org> wrote:
> # Not implemented formats. Ugly, but inevitable. This is the same as
> # issue bpo-2531: equality is also used for membership testing and must
> # return a result.
> a = array.array('u', 'xyz')
> v = memoryview(a)
> self.assertNotEqual(v, a)
> self.assertNotEqual(a, v)I don't understand: a buffer format is a format for the struct module,
or for the array module?It's like this: memoryview follows the current struct syntax, which
doesn't have 'u'. memory_richcompare() does not understand 'u', but
is required to return something for __eq__ and __ne__, so it returns
'not equal'.This isn't so important, since I discovered (see my later post)
that 'u' and 'w' were scheduled for inclusion in the struct
module anyway.So I think we should focus on whether the proposed 'c', 'u' and 'w'
format specifiers still make sense after the PEP-393 changes.@Stefan: What is the status of this issue?
I'm not sure what to do. Martin's opinion was that the change should
be reverted:http://mail.python.org/pipermail/python-dev/2012-March/117390.html
Should we do something before Python 3.3 final?
Is it possible without too much effort to keep the old behavior
('u' -> Py_UNICODE)? Then I'd say that should go into 3.3.The problem with the current behavior is that it's neither backwards
compatible nor PEP-3118 compliant.If it is too much work to restore the status quo, we could leave this
change with the excuse that 'u' is probably not used very often.Here is a patch reverting changes of the PEP-393, as suggested by Martin von Loewis. With the patch, array uses Py_UNICODE* type for the 'u' format. So array.array('u', '\u0010ffff')[0] should return '\uDBFF' on Windows.
The diff between b9558df8cc58 and default with array_revert_pep393.patch
applied is small, but I noticed that in some places you switched back to
Py_UNICODE typecode and in others not. For instance, in struct arraydescr
typecode is still char.I'm not sure why typecode was originally Py_UNICODE though.
The diff between b9558df8cc58 and default with array_revert_pep393.patch
applied is small, but I noticed that in some places you switched back to
Py_UNICODE typecode and in others not.I just copied code from Python 3.2, I forgot to update typecode type
(Py_UNICODE => char). I attach a new patch which changes also the
documentation of the "u" format.array_revert_pep393-2.patch looks good (checked against 7042a83f37e
and all following commits that should be kept).@georg: are you ok with this change? It reverts the behaviour of Python 3.2 and avoids to have to maintain an API that nobody wants to use ('u' format using Py_UCS4, 32 bits unsigned).
36 remaining items
Stefan, your patch array_deprecate_u.diff is fine. If you get to it, please also rephrase the clause "Python's unicode type"; not sure what the convention is to refer to Py_UNICODE now (perhaps "historical unicode type").
New changeset 9c7515e29219 by Stefan Krah in branch 'default':
Issue bpo-13072: The array module's 'u' format code is now deprecated and
http://hg.python.org/cpython/rev/9c7515e29219Good, I think this can be closed then.
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Aug 24, 2012 This has been deprecated for a long time, I was unable to find any uses in the top 1000 pypi projects or on gh. This should be set for removal in 3.15.
Its removal is scheduled for Python 3.16, see the array C code:
if (PyErr_WarnEx(PyExc_DeprecationWarning, "The 'u' type code is deprecated and " "will be removed in Python 3.16", 1)) { return NULL; } }I see now why I was mislead, the note is under
pending-removal-in-future-versionstoo. I sent a pr.- added a commit that references this issue
on Apr 29, 2025 - moved this to Todo in Struct, memoryview and array issues 🏗️
on Jul 13, 2026 - moved this from Todo to Done in Struct, memoryview and array issues 🏗️
on Jul 13, 2026
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs