Top AI Repos โ open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
๐ญ Mustard is a Swift library for tokenizing strings when splitting by whitespace doesn't cut it.
| Date | Stars |
|---|---|
| 2026-07-24 | 686 |
| 2026-07-25 | 686 |
| 2026-07-28 | 686 |
| 2026-07-30 | 686 |
| 2026-08-06 | 686 |
Today
โ stars today
This week
โ stars this week
This month
โ stars this month
Momentum
0.0
growth rate 0.00%/day
# Mustard ๐ญ
[](https://github.com/mathewsanders/Mustard/blob/master/LICENSE) [](https://github.com/Carthage/Carthage) [](https://swift.org/package-manager/)
Mustard is a Swift library for tokenizing strings when splitting by whitespace doesn't cut it.
## Quick start using character sets
Foundation includes the `String` method [`components(separatedBy:)`](https://developer.apple.com/documentation/foundation/nsstring/1413214-components) that allows us to get substrings divided up by certain characters:
````Swift
let sentence = "hello 2017 year"
let words = sentence.components(separatedBy: .whitespaces)
// words.count -> 3
// words = ["hello", "2017", "year"]
````
Mustard provides a similar feature, but with the opposite approach, where instead of matching by separators you can match by one or more character sets, which is useful if separators simply don't exist:
````Swift
import Mustard
let sentence = "hello2017year"
let words = sentence.components(matchedWith: .letters, .decimalDigits)
// words.count -> 3
// words = ["hello", "2017", "year"]
````
If you want more than just the substrings, you can use the `tokens(matchedWith: CharacterSet...)` method which will return an array of `TokenType`.
As a minimum, `TokenType` requires properties for text (the substring matched), and range (the range of the substring in the original string). When using CharacterSets as a tokenizer, the more specific type `CharacterSetToken` is returned, which includes the property `set` which contains the instance of CharacterSet that was used to create the match.
````Swift
import Mustard
let tokens = "123Hello world&^45.67".tokens(matchedWith: .decimalDigits, .letters)
// tokens: [CharacterSet.Token]
// tokens.count -> 5 (characters '&', '^', and '.' are ignored)
//
// second token..
// token[1].text -> "Hello"
// token[1].range -> Range<String.Index>(3..<8)
// token[1].set -> CharacterSet.letters
//
// last token..
// tokens[4].text -> "67"
// tokens[4].range -> Range<String.Index>(19..<21)
// tokens[4].set -> CharacterSet.decimalDigits
````
## Advanced matching with custom tokenizers
Mustard can do more than match from character sets. You can create your own tokenizers with more
sophisticated matching behavior by implementing the `TokenizerType` and `TokenType` protocols.
Here's an example of using `DateTokenizer` ([see example for implementation](Documentation/Template%20tokenizer.md)) that finds substrings that match a `MM/dd/yy` format.
`DateTokenizer` returns tokens with the type `DateToken`. Along with the substring text and range, `DateToken` includes a `Date` object corresponding to the date in the substring:
````Swift
import Mustard
let text = "Serial: #YF 1942-b 12/01/17 (Scanned) 12/03/17 (Arrived) ref: 99/99/99"
let tokens = text.tokens(matchedWith: DateTokenizer())
// tokens: [DateTokenizer.Token]
// tokens.count -> 2
// ('99/99/99' is *not* matched by `DateTokenizer` because it's not a valid date)
//
// first date
// tokens[0].text -> "12/01/17"
// tokens[0].date -> Date(2017-12-01 05:00:00 +0000)
//
// last date
// tokens[1].text -> "12/03/17"
// tokens[1].date -> Date(2017-12-03 05:00:00 +0000)
````
## Documentation & Examples
- [Greedy tokens and tokenizer order](Documentation/Greedy%20tokens%20and%20tokenizer%20order.md)
- [Token types and AnyToken](Documentation/Token%20types%20and%20AnyToken.md)
- [TokenizerType: implementing your own tokenizer](Documentation/TokenizerType%20protocol.md)
- [EmojiTokenizer: matching emoji substrings](Documentation/Matching%20emoji.md)
- [LiteralTokenizer: matching specific substrings](Documentation/Literal%20tokenizer.md)
- [DateTokenizer: tokenizer based on template match](DoExcerpt of 4,795 characters
Read on GitHub53
Muescha ยท Germany
1
1
Would you bet a product on this? Bounded 0โ100 and slow moving.
matched fp:dcd521000e41e4eb, topic:tokenizer, readme:tokenizer