Visitar URL original
3.12+: TokenError no longer raised for some invalid source (mismatched braces?) · Issue #105714 · python/cpython · GitHub
Skip to content

3.12+: TokenError no longer raised for some invalid source (mismatched braces?) #105714

Description

@asottile

this comes from the pycodestyle testsuite -- I expect the tokenizer to raise an error for this however it seems to silently accept it now.

Bug report

yep, this is just a single curly brace

}
$ python3.11 -m tokenize t.py
t.py:2:0: error: EOF in multi-line statement
$ python3.12 -m tokenize t.py
0,0-0,0:            ENCODING       'utf-8'        
1,0-1,1:            OP             '}'            
1,1-1,2:            NEWLINE        '\n'           
2,0-2,0:            ENDMARKER      '' 

Your environment

  • CPython versions tested on: d310fc7
  • Operating system and architecture: ubuntu 22.04 x86_64

Activity

  1. asottile commented on Jun 13, 2023

    @asottile
    ContributorAuthor

    this seems to be caused by the change in #105061

  2. lysnikolaou commented on Jun 13, 2023

    @lysnikolaou
    Member

    Unfortunately, this seems like one of those where we won't be able to preserve the previous behavior.

    Tokenizing invalid pieces of code has been undefined (although undocumented) since forever, but we now have a warning explicitly mentioning just that since 3.11. There's also many different stakeholders pushing in opposite directions (code quality tools towards this being preserved, inspect & IPython/sympy profit from this being able to tokenize correctly).

    Sorry we can't be more helpful than this.

  3. asottile commented on Jun 13, 2023

    @asottile
    ContributorAuthor

    @lysnikolaou reverting the patch above restores the behaviour -- I definitely think this is doable and easy

  4. lysnikolaou commented on Jun 13, 2023

    @lysnikolaou
    Member

    This is doable, yes. However, without this patch, multiple other things break, one of which is inspect, on which a lot of other things depend like hypothesis in the case of the issue that triggered this.

    Pleasing both sets of stakeholders would require the C tokenizer to recover from errors and then raise at the end, which would be a much bigger undertaking and would burden us with more maintenance than we're comfortable with.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

3.12only security fixes3.13only security fixestype-bugAn unexpected behavior, bug, or error

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions