Skip to content

firstIndexOf, lastIndexOf and indices return UTF-16 offsets, which no other String function accepts #148

Description

@gilramir

Found against: gren 0.6.6, gren-lang/core 7.4.2, node 22

Summary

There's a combination of counting Unicode code points and UTF-16 code units in String,
without enough functions to cover all the bases.
Given this:

s = "𝄞ab"                      -- U+1D11E, then "a", then "b"

String.count s                 -- 3          (codepoints)
String.unitLength s            -- 4          (code units)

The MUSICAL SYMBOL G CLEF is a single Unicode code point (U+1D11E), but when encoded in UTF-16, it occupies 2 16-bit code units. (0xD834 0xDD1E).

"𝄞ab" is 3 Unicode code points (𝄞, a, and b)
But it has 4 "UTF-16 code units": 0xd834, 0xdd1e, 0x0061, 0x0062

String is character-oriented: count, slice, takeFirst, dropFirst,
reverse, toArray and pad all work in codepoints, and
#130 was fixed by making
_String_slice count codepoints rather than UTF-16 code units.

The three functions that produce an index were not part of that fix.
firstIndexOf, lastIndexOf and indices are String.prototype.indexOf and
lastIndexOf, so they answer in UTF-16 code units — the unit no other
function in the module accepts. For any string with a non-BMP character (i.e., code point >= U+10000) before
the match, an index handed straight back to slice points past where it was
found:

firstIndex returns a count of code units, but slice accepts a count of code poitns

s = "𝄞ab"                      -- U+1D11E, then "a", then "b"

String.firstIndexOf "a" s      -- Just 2     (code units; the character is at 1)
String.slice 2 3 s             -- "b"        (codepoints; "a" is at 1)

So the natural composition of firstIndexOf and slice is wrong for strings containing
astral characters (code points >= U+10000) — emoji, mathematical alphanumerics, most historic scripts, the
higher CJK extensions:

when String.firstIndexOf needle haystack is
    Just i ->
        String.slice i (i + String.count needle) haystack   -- not the match
    Nothing ->
        ""

String.indices has the same problem and compounds it: its loop advances by
sub.length, which is also code units, so every offset it returns is shifted by
the number of astral characters before it.

Reproduction

𝄞 is U+1D11E: one codepoint, two UTF-16 units. So in "𝄞ab" the character
"a" is at codepoint index 1 and code-unit index 2, and that one-unit
difference is the problem. needleAt below is the natural
composition — find a needle, slice at what was found — and it is the row to
read first.

module Main exposing (probe)


clef : String
clef =
    "𝄞ab"


needleAt : String -> String
needleAt needle =
    when String.firstIndexOf needle clef is
        Just i ->
            String.slice i (i + String.count needle) clef

        Nothing ->
            "not found"


showIndex : Maybe Int -> String
showIndex m =
    when m is
        Just n ->
            String.fromInt n

        Nothing ->
            "-"


probe : Array String
probe =
    [ needleAt "a"
    , needleAt "b"
    , needleAt "𝄞"
    , showIndex (String.firstIndexOf "a" clef)
    , showIndex (String.lastIndexOf "b" clef)
    , String.join "," (Array.map String.fromInt (String.indices "b" clef))
    , String.join "," (Array.map String.fromInt (String.indices "𝄞" "𝄞a𝄞"))
    , String.slice 1 2 clef
    , String.fromInt (String.count clef)
    , String.fromInt (String.unitLength clef)
    ]
expression expected, in codepoints actual
needleAt "a" "a" "b"
needleAt "b" "b" "" (index 3 is past the end)
needleAt "𝄞" "𝄞" "𝄞" (index 0, so nothing to shift)
String.firstIndexOf "a" clef Just 1 Just 2
String.lastIndexOf "b" clef Just 2 Just 3
String.indices "b" clef [2] [3]
String.indices "𝄞" "𝄞a𝄞" [0,2] [0,3]
String.slice 1 2 clef "a" "a" (consumers are already codepoints)
String.count clef 3 3
String.unitLength clef 4 4 (the one function specified in units)

The first two rows are the failure a caller sees: needleAt "a" returns the
character after the one it found, and needleAt "b" returns the empty string
because the index it was handed is past the end of a three-codepoint string.
Neither reports an error. A string of BMP-only text produces the expected
column throughout, which is why this has gone unnoticed.

Notes

The documentation does not currently say which unit any of the three answers
in. "Unless otherwise noted, all functions in this module deal with code
points" is the module's header, and these three are not noted.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions