FIX: Calculate accurate UTF-8 byte length for content-length in HTTPTarget - #3008
Open
Vardhman Gupta (Kaap10) wants to merge 1 commit into
Open
Vardhman Gupta (Kaap10) wants to merge 1 commit into
Vardhman Gupta (Kaap10) wants to merge 1 commit into
Conversation
…length in HTTPTarget
Contributor
Author
|
This is ready for review. Thanks! |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
While testing
HTTPTargetwith non-ASCII and multilingual prompts (such as Chinese characters and emojis), I noticed that requests were getting truncated on the receiving server.Looking into
HTTPTarget._send_prompt_to_target_async, thecontent-lengthheader was being calculated usinglen(http_body). In Python,len()on a string returns the character count instead of the actual encoded byte count. Per RFC 9110,Content-Lengthmust be the number of octets/bytes. For multi-byte characters (e.g."你好"which is 2 characters but 6 UTF-8 bytes), sendingcontent-length: 2caused the payload to be cut short.Summary of Changes:
_send_prompt_to_target_asyncto uselen(http_body.encode("utf-8"))for string bodies so the byte count is accurate.bytes/bytearraybodies.content-lengthwhenhttp_bodyis empty (e.g., bodylessGETrequests)."content-length"to keep dictionary keys consistent withparse_raw_http_request.Tests and Documentation
test_send_prompt_async_content_length_utf8_bytesto verify multi-byte characters calculate the correct byte length.test_send_prompt_async_content_length_recalculatedto check header updates when overriding raw requests.test_send_prompt_async_content_length_raw_bytesfor raw byte payload testing.test_http_target.pymock assertions.pytest tests/unit/prompt_target/(all passed) andruff check.