Kannada TN - Created Ordinal Semiotic Class - #493
Conversation
Signed-off-by: richa-2002 <richa@nvidia.com>
Signed-off-by: richa-2002 <richa@nvidia.com>
for more information, see https://pre-commit.ci
There was a problem hiding this comment.
The thing to settle is in exceptions.tsv: 1st–15th get a transliterated English reading while 16th and up get a native Kannada one, so 15th → ಫಿಫ್ಟೀನ್ತ್ but 16th → ಹದಿನಾರನೆಯ. Your own test file encodes both conventions, so it needs a decision rather than a bug fix. Separately, en_suffixes.tsv accepts any digit with any suffix (5st, 11st), which is worth constraining.
| th ನೆಯ | ||
| st ನೆಯ | ||
| nd ನೆಯ | ||
| rd ನೆಯ No newline at end of file |
There was a problem hiding this comment.
should-fix · verified
These four map unconditionally, with nothing tying a suffix to the digit it may follow — so the
grammar accepts English ordinals that don't exist:
5st -> ಐದನೆಯ
2th -> ಎರಡನೆಯ
7nd -> ಏಳನೆಯ
4rd -> ನಾಲ್ಕನೆಯ
11st -> ಹನ್ನೊಂದನೆಯ (English is 11th, never 11st)
In English the suffix is determined by the final digit, with 11/12/13 as the exception: st after
…1 (not 11), nd after …2 (not 12), rd after …3 (not 13), th everywhere else. Right now any
digit + any suffix normalizes silently, which means a typo or an unrelated token gets swallowed by
the ordinal class instead of falling through to word.
I haven't written a patch because the constraint is fiddly enough that an untested one would be
worse than none — but the repo idiom is to encode it structurally with pynini.difference on the
last digit rather than by weight, the way hi/taggers/roman.py gates its input set. Worth a test
case for at least 11th vs 11st once it's in.
There was a problem hiding this comment.
Removed en_suffixes.tsv and the English ordinal suffix handling, as it was allowing incorrect suffix combinations. The ordinal implementation now supports only Kannada ordinal forms.
| 1st ಫಸ್ಟ್ | ||
| 2nd ಸೆಕೆಂಡ್ | ||
| 3rd ಥರ್ಡ್ | ||
| 4th ಫೋರ್ಥ್ | ||
| 5th ಫಿಫ್ತ್ | ||
| 6th ಸಿಕ್ಸ್ತ್ | ||
| 7th ಸೆವೆಂತ್ | ||
| 8th ಎಯ್ತ್ | ||
| 9th ನೈನ್ತ್ | ||
| 10th ಟೆನ್ತ್ | ||
| 11th ಇಲೆವೆಂತ್ | ||
| 12th ಟ್ವೆಲ್ಫ್ತ್ | ||
| 13th ಥರ್ಟೀನ್ತ್ | ||
| 14th ಫೋರ್ಟೀನ್ತ್ | ||
| 15th ಫಿಫ್ಟೀನ್ತ್ |
There was a problem hiding this comment.
should-fix · verified + linguistic-claim
This block gives 1st–15th a transliterated English reading, but anything from 16 up falls
through to the compositional path and gets a native Kannada reading. So the same orthographic
pattern flips reading systems mid-range:
14th -> ಫೋರ್ಟೀನ್ತ್ (English)
15th -> ಫಿಫ್ಟೀನ್ತ್ (English)
16th -> ಹದಿನಾರನೆಯ (Kannada) <- discontinuity
17th -> ಹದಿನೇಳನೆಯ (Kannada)
Same for Kannada digits with an English suffix (೧೫th -> ಫಿಫ್ಟೀನ್ತ್, ೧೬th -> ಹದಿನಾರನೆಯ), and the
split shows up in st/nd/rd too — 1st -> ಫಸ್ಟ್ but 21st -> ಇಪ್ಪತ್ತೊಂದನೆಯ.
Your own test file encodes both sides of this: L37-38 assert ೮೮th~ಎಂಬತ್ತೆಂಟನೆಯ / 88th~ಎಂಬತ್ತೆಂಟನೆಯ
(Kannada), while this file asserts 15th~ಫಿಫ್ಟೀನ್ತ್ (English). So it isn't a bug I'm inferring —
it's two conventions coexisting.
I don't speak Kannada, so I can't say which is right, and both are defensible: reading Latin-suffixed
ordinals as transliterated English is reasonable for code-switched text, and so is always reading the
number natively. What isn't defensible is the boundary at 15. Could you pick one and make it uniform?
- If English for Latin suffixes: the exception list needs to keep going (16th, 20th, 21st, 100th…),
which argues for generating it rather than enumerating. - If Kannada everywhere: drop the
1st–15throws here and keep only the genuinely irregular
1ನೇ/೧ನೇ-> ಮೊದಲನೆಯ at L1-2.
There was a problem hiding this comment.
Made the ordinal handling uniform by selecting the Kannada convention everywhere. Removed the English transliterated ordinal readings and kept Kannada ordinal forms consistently.
| ು | ||
| ಿ No newline at end of file |
There was a problem hiding this comment.
optional · verified + linguistic-claim
Nice use of cdrewrite to strip the cardinal's final vowel before the suffix — and the two entries
are well chosen: ು covers the digits/teens/ties and ಿ covers the three scale_suffixes.tsv rows
ending that way (ಕೋಟಿ). I checked and every cardinal in data/numbers/ ends in one of the two.
Except zero: ಸೊನ್ನೆ ends in ೆ, which isn't here, so nothing gets stripped:
0ನೇ -> ಸೊನ್ನೆನೆಯ
೦ನೇ -> ಸೊನ್ನೆನೆಯ
Whether that even matters is your call — "zeroth" is a marginal thing to write, and rejecting 0ನೇ
outright (so it falls through to word) may well be better than normalizing it. But right now it
produces output, and I'd rather flag that it's unreviewed than let it through silently. If ಸೊನ್ನೆನೆಯ
isn't the form a Kannada reader expects, either add ೆ here or exclude zero from the ordinal input
set.
There was a problem hiding this comment.
Added 0ನೇ and ೦ನೇ as exceptions in exceptions.tsv with the output ಸೊನ್ನೆಯ, since zero does not follow the existing last-vowel stripping pattern.
| 21st~ಇಪ್ಪತ್ತೊಂದನೆಯ | ||
| ಈ ಮಗುವಿನ ಶೈಕ್ಷಣಿಕ ಫಲಿತಾಂಶವು ಯಾವಾಗಲೂ ಇಡೀ ತರಗತಿಯಲ್ಲಿ ೧ನೇ ಸ್ಥಾನದಲ್ಲಿದೆ~ಈ ಮಗುವಿನ ಶೈಕ್ಷಣಿಕ ಫಲಿತಾಂಶವು ಯಾವಾಗಲೂ ಇಡೀ ತರಗತಿಯಲ್ಲಿ ಮೊದಲನೆಯ ಸ್ಥಾನದಲ್ಲಿದೆ | ||
| ಈ ಸಾಲಿನಿಂದ ಕೆಳಕ್ಕೆ ಎಣಿಸಿದಾಗ 5ನೇ ಶಿಯಾವೋ ಮಿಂಗ್.~ಈ ಸಾಲಿನಿಂದ ಕೆಳಕ್ಕೆ ಎಣಿಸಿದಾಗ ಐದನೆಯ ಶಿಯಾವೋ ಮಿಂಗ್. | ||
| ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ೧೦೧ನೇ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ.~ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ನೂರ ಒಂದನೆಯ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ. |
There was a problem hiding this comment.
should-fix · verified
81 cases with both scripts, tier boundaries, a 12-digit number and two full sentences — this is
genuinely good coverage, and the whole suite is green (424 passed on a fresh cache). Two gaps, both
of which hide things I've commented on elsewhere:
1. Nothing covers the 15/16 reading boundary. 88th is here but no case sits either side of the
discontinuity, so whichever convention you settle on in exceptions.tsv, these pin it down. Every
expected value below is copied from an actual run against this PR:
| ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ೧೦೧ನೇ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ.~ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ನೂರ ಒಂದನೆಯ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ. | |
| ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ೧೦೧ನೇ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ.~ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ನೂರ ಒಂದನೆಯ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ. | |
| 16th~ಹದಿನಾರನೆಯ | |
| 21st~ಇಪ್ಪತ್ತೊಂದನೆಯ | |
| ೧೦೦೦ನೇ~ಒಂದು ಸಾವಿರನೆಯ |
I've left 15th out on purpose — what it should produce depends on which convention you pick, and
guessing at it would be worse than leaving it to you. (I nearly shipped a guessed thousand-tier form
here too; the real output turned out to be ಒಂದು ಸಾವಿರನೆಯ, not what I'd assumed, which is why it's above verbatim from a run.)
2. Nothing covers malformed English suffixes, which is why 5st and 11st currently normalize
unnoticed. Once en_suffixes.tsv is constrained, a guard belongs here — 11th normalizes today
(ಇಲೆವೆಂತ್) while 11st should ideally be left alone.
There was a problem hiding this comment.
Added the missing boundary cases to the test file, including 15ನೇ, 16ನೇ using the Kannada convention consistently.
Signed-off-by: richa-2002 <richa@nvidia.com>
for more information, see https://pre-commit.ci
|
Thank you for the review @folivoramanh , I've addressed all the changes u suggested. Testing Pytest: All test cases passed — Ordinal. |
folivoramanh
left a comment
There was a problem hiding this comment.
Re-reviewed at c6b8fe5b. The previous round's fixes check out: en_suffixes.tsv is gone with no dangling references, the English transliterations are out of exceptions.tsv, 0ನೇ/೦ನೇ → ಸೊನ್ನೆಯ both work, 15ನೇ was added, and 88th/21st now correctly fall through to word — 414/414 pytest pass on a fresh cache and the TN export compiles clean. The one blocker left is that both exceptions are keyed only on the ನೇ spelling, so the ನೆ variant your own test file uses elsewhere still produces ಸೊನ್ನೆನೆಯ for zero and ಒಂದನೆಯ (not ಮೊದಲನೆಯ) for one. Sparrowhawk was not executed here (no Docker) — please paste your --MODE=test output.
| ೦ನೇ ಸೊನ್ನೆಯ | ||
| 0ನೇ ಸೊನ್ನೆಯ No newline at end of file |
There was a problem hiding this comment.
blocker · verified
The exception table is keyed only on the ನೇ spelling, but kn_suffixes.tsv accepts ನೆ as well — and your own test file uses that variant for regular numbers (೧೭೮೨ನೆ, 6789876ನೆ, 10000000000000ನೆ), so it is valid input by this PR's own definition. Both exceptions therefore fall through to the general path on the short variant. I ran all four at c6b8fe5b:
0ನೆ→ಸೊನ್ನೆನೆಯ,೦ನೆ→ಸೊನ್ನೆನೆಯ— the exact double-suffix artifact this round removed forನೇ(expectedಸೊನ್ನೆಯ)1ನೆ→ಒಂದನೆಯ,೧ನೆ→ಒಂದನೆಯ— not theಮೊದಲನೆಯyour table declares for೧ನೇ
Whichever reading is right for ೧ನೆ, the two spellings shouldn't disagree.
| ೦ನೇ ಸೊನ್ನೆಯ | |
| 0ನೇ ಸೊನ್ನೆಯ | |
| ೦ನೇ ಸೊನ್ನೆಯ | |
| 0ನೇ ಸೊನ್ನೆಯ | |
| 1ನೆ ಮೊದಲನೆಯ | |
| ೧ನೆ ಮೊದಲನೆಯ | |
| ೦ನೆ ಸೊನ್ನೆಯ | |
| 0ನೆ ಸೊನ್ನೆಯ |
If you'd rather not double every future exception row, the structural alternative is to key exceptions.tsv on the digit alone and let kn_suffixes consume the suffix generically in ordinal.py:41-49 — same effect, no combinatorial growth.
| ಈ ಸಾಲಿನಿಂದ ಕೆಳಕ್ಕೆ ಎಣಿಸಿದಾಗ 5ನೇ ರಾಘವೇಂದ್ರ.~ಈ ಸಾಲಿನಿಂದ ಕೆಳಕ್ಕೆ ಎಣಿಸಿದಾಗ ಐದನೆಯ ರಾಘವೇಂದ್ರ. | ||
| ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ೧೦೧ನೇ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ.~ಅಭಿನಂದನೆಗಳು! ನೀವು ನಮ್ಮ ಅಂಗಡಿಯ ನೂರ ಒಂದನೆಯ ಗ್ರಾಹಕರಾಗಿದ್ದೀರಿ. | ||
| ಈ ಪಟ್ಟಿಯ ಆರಂಭದಿಂದ ೧೦೦ನೇ ವ್ಯಕ್ತಿಯವರೆಗೆ ಇರುವ ಎಲ್ಲರೂ ನಿಮ್ಮ ಗುರಿ ಗ್ರಾಹಕರು.~ಈ ಪಟ್ಟಿಯ ಆರಂಭದಿಂದ ನೂರನೆಯ ವ್ಯಕ್ತಿಯವರೆಗೆ ಇರುವ ಎಲ್ಲರೂ ನಿಮ್ಮ ಗುರಿ ಗ್ರಾಹಕರು. | ||
| 0ನೇ~ಸೊನ್ನೆಯ No newline at end of file |
There was a problem hiding this comment.
should-fix · verified
Only 0ನೇ is covered here. ೦ನೇ is in exceptions.tsv but untested (the repo convention is to test native-digit values in both scripts), and no exception number is tested with the ನೆ variant — which is why the bug on exceptions.tsv above passes CI. These lines fail today and pass once that fix lands:
| 0ನೇ~ಸೊನ್ನೆಯ | |
| 0ನೇ~ಸೊನ್ನೆಯ | |
| ೦ನೇ~ಸೊನ್ನೆಯ | |
| 0ನೆ~ಸೊನ್ನೆಯ | |
| ೦ನೆ~ಸೊನ್ನೆಯ | |
| 1ನೆ~ಮೊದಲನೆಯ | |
| ೧ನೆ~ಮೊದಲನೆಯ |
Signed-off-by: richa-2002 <richa@nvidia.com>
|
Thank you for the review @folivoramanh , I've addressed all the changes u suggested. Testing Pytest: All test cases passed — Ordinal. GRAMMARS = tn_grammars Ran 4 tests. OK |
What does this PR do ?
Add a one line overview of what this PR aims to accomplish.
Before your PR is "Ready for review"
Pre checks:
git commit -sto sign.pytestor (if your machine does not have GPU)pytest --cpufrom the root folder (given you marked your test cases accordingly@pytest.mark.run_only_on('CPU')).bash tools/text_processing_deployment/export_grammars.sh --MODE=test ...pytestand Sparrowhawk here.__init__.pyfor every folder and subfolder, includingdatafolder which has .TSV files?Copyright (c) 2023, NVIDIA CORPORATION & AFFILIATES. All rights reserved.to all newly added Python files?Copyright 2015 and onwards Google, Inc.. See an example here.try import: ... except: ...) if not already done.PR Type:
If you haven't finished some of the above items you can still open "Draft" PR.