Skip to content

Fix(CMEM-8068): preserve non-ASCII characters in JSON output - #10

Merged
msaipraneeth merged 7 commits into
mainfrom
bugfix/UnicodeEscape-CMEM-8068
Sep 8, 2026
Merged

msaipraneeth merged 7 commits into
mainfrom
bugfix/UnicodeEscape-CMEM-8068

Conversation

@msaipraneeth

Copy link
Copy Markdown
Contributor

json.dumps() defaults to ensure_ascii=True, which escaped special characters (e.g. ö, ü) into unicode escape sequences both in entity values built from nested GraphQL response data (create_entity) and in the JSON payload written to the target dataset.

json.dumps() defaults to ensure_ascii=True, which escaped special
characters (e.g. ö, ü) into unicode escape sequences both in entity
values built from nested GraphQL response data (create_entity) and in
the JSON payload written to the target dataset.
@github-actions

github-actions Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Coverage

Coverage Report
File Stmts Miss Cover Missing
init.py 0 0 100%
workflow/init.py 0 0 100%
workflow/graphql.py 107 5 95% 163 179-180 187 208
workflow/utils.py 15 0 100%
TOTAL 122 5 96%  

Tests Skipped Failures Errors Time
14 0 💤 0 ❌ 0 🔥 23.575 ⏱

The public test endpoint's fruit id 4 (Limón / Rutáceae) genuinely
contains non-ASCII characters, so it can drive a real end-to-end test
of the write_to_dataset() JSON path instead of only the isolated
create_entity() unit test. Checks the raw file bytes (not just the
json.loads()'d result, which would hide the bug either way) for the
absence of the escape sequences.
Neither function is called from GraphQLPlugin.execute() or anywhere
else in the package - the plugin builds entities via cmem-plugin-base's
build_entities_from_data() instead. Confirmed no external callers.
The @plugin documentation previously said nothing about the ports, the
Jinja-templating value shapes, or what happens when Query/Query
variables contains Jinja syntax but no input is connected (a crash for
Query, a silent no-op for Query variables). Also corrected the earlier
CHANGELOG Fixed entry, which still described the entity-values path
that turned out to be dead code and was removed.
cmem_client ships no py.typed marker, so mypy treats
client.files.read() as returning Any regardless of its actual bytes
return type. An explicit bytes annotation on the intermediate variable
lets mypy trust it, so .decode("utf-8") resolves to str as declared.
write_to_dataset()/post_file_resource() are documented and typed to
take a byte stream, but the target-dataset write passed an
io.StringIO text stream instead. This masqueraded as working while
ensure_ascii=True guaranteed pure-ASCII output (1 char == 1 byte); now
that real UTF-8 multi-byte characters (e.g. from the ensure_ascii=False
fix) flow through, the byte/char-length mismatch silently truncated
the uploaded file, as seen by CI's non-ASCII regression test failing
with a JSONDecodeError on a file missing its closing bracket.

Switched to io.BytesIO(...encode("utf-8")), matching the pattern
already used correctly elsewhere (e.g. cmem-plugin-jq,
cmem-plugin-validation).
The example line kept the docstring's source indentation, which Markdown
renders as an indented code block instead of flowing text.

Claude-Session: https://claude.ai/code/session_01XE4Ptbbj45VPRQnPsR9Lvm
@msaipraneeth
msaipraneeth merged commit cd16578 into main Sep 8, 2026
2 checks passed
@msaipraneeth
msaipraneeth deleted the bugfix/UnicodeEscape-CMEM-8068 branch September 8, 2026 11:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant