prithivMLmods/Coder-Stat
Coder-Stat Dataset Overview The Coder-Stat dataset is a collection of programming-related data, including problem IDs, programming languages, original statuses, and source code snippets. This dataset is designed to assist in the analysis of coding patterns, error types, and performance metrics. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text: Contains text data, including source code snippets.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Coder-Stat.
3139
1 2<h1>3Kanglish : Analysis on Artificial Language 4</h1>5 6<p>The late Prof. Kanazawa made an artificial language named Kanglish, which is 7similar to English, for studying mythology. Words and sentences of Kanglish are 8written with its own special characters called "Kan-characters". The size of the 9set of the Kan-characters is 38, i.e., there are 38 different Kan-characters in 10the set. Since Kan-characters cannot be directly stored in a computer because of 11the lack of a coded character set, Prof. Kanazawa devised a way to represent 12each Kan-character as an alphabetical letter or an ordered combination of two 13alphabetical letters. Thus, each Kan-character is represented as one of the 14following 26 letters 15</p><blockquote>"a", "b", "c", "d", "e", "f", "g", "h", "i", "j", "k", "l", "m", 16 "n", "o", "p", "q", "r", "s", "t", "u", "v", "w", "x", "y", and "z", 17</blockquote>or one of the following 12 combinations of letters 18<blockquote>"ld", "mb", "mp", "nc", "nd", "ng", "nt", "nw", "ps", "qu", "cw", 19 and "ts". </blockquote>20<p>21In addition, the Kan-characters are ordered according to 22the above alphabetical representation. The order is named Kan-order in which the 23Kan-character represented by "a" is the first one, that by "b" is the second, 24that by "z" is the 26th, that by "ld" is the 27th, and that by "ts" is the 38th 25(the last). 26</p>27<p></p>28<p>The representations of words in Kanglish are separated by spaces. Each 29sentence is written in one line, so there are no periods or full stops in 30Kanglish. All alphabetical letters are in lower case, i.e., there are no 31capitalizations. </p>32<p>We currently have many documents written in Kanglish with the alphabetical 33representation. However, we have lost most of Prof. Kanazawa's work on how they 34can be translated. To recognize his work, we have decided to first analyze them 35statistically. The first analysis is to check sequences of consecutive 36Kan-characters in words. </p>37<p>For example, a substring "ic" in a word "quice" indicates an ordered pair of 38two adjacent Kan-characters that are represented by "i" and "c". For simplicity, 39we make a rule that, in the alphabetical representation of a word, a 40Kan-character is recognized as the longest possible alphabetical representation 41from left to right. Thus, a substring "ncw" should be considered as a pair of 42"nc" and "w". It does not consist of "n" and "cw", nor "n", "c", and "w". </p>43<p>For each Kan-character, there are 38 possible pairs of that Kan-character and 44another Kan-character, e.g. "aa", "ab", ..., "az", "ald", ..., "ats". Thus, 45mathematically, there is a total of 1444 (i.e., 38x38) possible pairs, including 46pairs such as "n" and "cw", which is actually not allowed according to the above 47rule. </p>48<p>Your job is to write a program that counts how many times each pair occurs in 49input data. For example, in the sentence </p><pre> qua ist qda quang quice</pre>50<p>51the Kan-character represented by "qu" appears 52three times. There are two occurrences of the pair of "qu" and "a", and one 53occurrence of the pair of "qu" and "i". Note that the alphabetical letter "q" 54appears four times in the example sentence, but the Kan-character represented by 55"q" occurs only once, because "qu" represents another Kan-character that is 56different from the Kan-character represented by "q". 57<p></p>58<p>For simplicity, a newline at the end of a line is considered as a space. Thus 59in the above example, "e" is followed by a space. </p>60<h2>Input</h2><pre><i>n</i>61<i>line</i><sub>1</sub>62<i>line</i><sub>2</sub>63...64<i>line</i><sub><i>n</i></sub>65</pre>66<p>The first line of the input is an integer <i>n</i>, which indicates the 67number of lines that follow. Each line except for the first line represents one 68Kanglish sentence. You may assume that <i>n</i> <= 1000 and that each line 69has at most 59 alphabetical letters including spaces. </p>70<h2>Output</h2><pre>a <i>kc</i><sub>1</sub> <i>m</i><sub>1</sub>71b <i>kc</i><sub>2</sub> <i>m</i><sub>2</sub>72c <i>kc</i><sub>3</sub> <i>m</i><sub>3</sub>73...74ts <i>kc</i><sub>38</sub> <i>m</i><sub>38</sub>75</pre>76<p>The output consists of 38 lines for the whole input lines. Each line of the 77output has two strings and an integer. In the <i>i</i>-th line in the output, 78the first string is the alphabetical representation of the <i>i</i>-th 79Kan-character in the Kan-order. For example, the first string of the first line 80is "a", that of the third line is "c", and that of the 37th line is "cw". The 81first string is followed by a space. </p>82<p>The second string in the <i>i</i>-th line (denoted by <i>kc<sub>i</sub></i> 83above) shows the alphabetical representation of a Kan-character that most often 84occurred directly after the first Kan-character. If there are two or more such 85Kan-characters, the first one in the Kan-order should be printed. The second 86string is followed by a space. </p>87<p>The integer (denoted by <i>m<sub>i</sub></i> above) in the <i>i</i>-th line 88shows the number of times that the second Kan-character occurred directly after 89the first Kan-character. In other words, the integer shows the number of times 90that ``the ordered pair of the first Kan-character and the second 91Kan-character'' appeared in the input. The integer is followed by a newline. 92</p>93<p>Suppose the 28th output line is as follows: 94</p><blockquote><pre>mb e 4</pre></blockquote>95<p>96"mb" is output because it is the 28th character in 97the Kanglish alphabet. "e 4" means that the pair "mbe" appeared 4 times in the 98input, and that there were no pairs beginning with "mb" that appeared more than 994 times. 100<p></p>101<p>Note that if the <i>i</i>-th Kan-character does not appear in the input, or 102if the <i>i</i>-th Kan-character is not followed by any other Kan-characters but 103spaces, the second string in the <i>i</i>-th output line should be "a" and the 104third item should be zero. </p>105<p>Although the output does not include spaces, Kan-characters that appear with 106a space in-between is not considered as a pair. Thus, in the following example 107</p><blockquote><pre>abc def</pre></blockquote>108<p>109"d" is not counted as occurring after "c". 110<p></p>111<h2>Sample Input</h2>112<pre>1133114nai tiruvantel ar varyuvantel i valar tielyama nu vilya115qua ist qda quang ncw psts116svampti tsuldya jay quadal ciszeriol117</pre>118<h2>Output for the Sample Input</h2>119<pre>120a r 3121b a 0122c i 1123d a 2124e l 3125f a 0126g a 0127h a 0128i s 2129j a 1130k a 0131l y 2132m a 1133n a 1134o l 1135p a 0136q d 1137r i 1138s t 1139t i 3140u v 2141v a 5142w a 0143x a 0144y a 3145z e 1146ld y 1147mb a 0148mp t 1149nc w 1150nd a 0151ng a 0152nt e 2153nw a 0154ps ts 1155qu a 3156cw a 0157ts u 1158</pre>159 160 