Fix(CMEM-8068): preserve non-ASCII characters in JSON output - #10
Merged
Merged
Conversation
json.dumps() defaults to ensure_ascii=True, which escaped special characters (e.g. ö, ü) into unicode escape sequences both in entity values built from nested GraphQL response data (create_entity) and in the JSON payload written to the target dataset.
Coverage Report
|
The public test endpoint's fruit id 4 (Limón / Rutáceae) genuinely contains non-ASCII characters, so it can drive a real end-to-end test of the write_to_dataset() JSON path instead of only the isolated create_entity() unit test. Checks the raw file bytes (not just the json.loads()'d result, which would hide the bug either way) for the absence of the escape sequences.
Neither function is called from GraphQLPlugin.execute() or anywhere else in the package - the plugin builds entities via cmem-plugin-base's build_entities_from_data() instead. Confirmed no external callers.
The @plugin documentation previously said nothing about the ports, the Jinja-templating value shapes, or what happens when Query/Query variables contains Jinja syntax but no input is connected (a crash for Query, a silent no-op for Query variables). Also corrected the earlier CHANGELOG Fixed entry, which still described the entity-values path that turned out to be dead code and was removed.
cmem_client ships no py.typed marker, so mypy treats
client.files.read() as returning Any regardless of its actual bytes
return type. An explicit bytes annotation on the intermediate variable
lets mypy trust it, so .decode("utf-8") resolves to str as declared.
write_to_dataset()/post_file_resource() are documented and typed to
take a byte stream, but the target-dataset write passed an
io.StringIO text stream instead. This masqueraded as working while
ensure_ascii=True guaranteed pure-ASCII output (1 char == 1 byte); now
that real UTF-8 multi-byte characters (e.g. from the ensure_ascii=False
fix) flow through, the byte/char-length mismatch silently truncated
the uploaded file, as seen by CI's non-ASCII regression test failing
with a JSONDecodeError on a file missing its closing bracket.
Switched to io.BytesIO(...encode("utf-8")), matching the pattern
already used correctly elsewhere (e.g. cmem-plugin-jq,
cmem-plugin-validation).
The example line kept the docstring's source indentation, which Markdown renders as an indented code block instead of flowing text. Claude-Session: https://claude.ai/code/session_01XE4Ptbbj45VPRQnPsR9Lvm
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
json.dumps() defaults to ensure_ascii=True, which escaped special characters (e.g. ö, ü) into unicode escape sequences both in entity values built from nested GraphQL response data (create_entity) and in the JSON payload written to the target dataset.