One of the manga series I've been waiting forever for an English translation of (to buy, that is; fan-translated versions have been around online for years) is Elfen Lied, which I mentioned being fond of in the past. Well, there's still not official English translation, but I just learned something very interesting: there's an official Spanish version (also German and Taiwanese Chinese, for speakers of those languages). So I can own the series I've wanted for a long time, while practicing my Spanish.
I mentioned a ways back that I was reading the Spanish version of the Trique grammar book (as there wasn't an equivalent English version), and I was surprised how well I could read the Spanish, without even looking things up. However, a part of the reason this went so smoothly was that most of the unknown words in that were linguistic jargon that I was able to figure out from context very quickly.
Actually, my very first exposure to Trique, prior to that, also involved Spanish. It all began with a Spanish/Trique bilingual Bible I got from my grandpa. I started by comparing the text in Spanish and Trique to attempt to figure out the grammar by example. This was somewhat more difficult than the grammar book. It had a much broader array of words used, so it was significantly harder to figure out unknown words from context. Still, I made a respectable amount of progress, given the method.
If I'm lucky, reading Elfen Lied would be closer to the former, as the illustrations give some context. But even if I infrequently need to look things up it wouldn't be too bad. And it would very likely improve my ability to read/write Spanish in general by a noteworthy amount.
So, I'll probably buy that after money is no longer so tight.
Search This Blog
Showing posts with label linguistics. Show all posts
Showing posts with label linguistics. Show all posts
Monday, March 22, 2010
Sunday, March 21, 2010
Q's Squishy Thing of the Day
Just saw this today. Perhaps nobody else will find it interesting, but I do.
Essential WoW terminology in other languages
Essential WoW terminology in other languages
Labels:
linguistics,
randomthoughts,
squishy
Wednesday, July 15, 2009
& the Genitive
Last in the series of core cases (the four cases almost universal among languages that use grammatical case) is the genitive (the nominative/absolutive and accusative/ergative are described here, and the dative here).
Genitive Relation
The genitive relation consists of a noun phrase (the minimum noun phrase in most languages being a single noun) that modifies another noun phrase. As described by grammar books, a genitive phrase is a noun phrase that specifies or describes another noun phrase, with the result being that the genitive phrase helps to answer the question "which?" with regard to the modified noun phrase; e.g. "the girl next door's dog" (genitive in italic) answers "which dog?", "music of Starcraft" answers "which music?", etc.
As it essentially contains all manner of modifying phrases, clearly the genitive is an extremely broad relation, and there are quite a few more-specific relations that fall under the umbrella of "genitive". A semi-exhaustive listing of the various types of relations the genitive can express:
One very good reason for a language to have more than a single mechanism of representing the genitive is that any single such mechanism would necessarily be ambiguous. Take, for example, the English phrase "betrayal of Illidan"; with this phrase it is unclear whether the genitive is subjective or objective - that is, whether Illidan is the betrayer or the one betrayed. Other such examples can be invented based on other uses of the genitive.
Of course it's not all bad; if there were no ambiguity in language, puns would be impossible. One real-life example of genitive ambiguity comes from IRC. jfroy says 'Snow Leopard meeting' in the sense of 'meeting about [Apple] Snow Leopard', to which I reply, interpreting that as 'meeting with a snow leopard':
<@jfroy> Snow Leopard meeting over. Ow.
<enma_hinobara>Next time don't poke it
<enma_hinobara>It's not Kaity
The Genitive in English
It's arguable whether English still has a genitive case; if it does, it only exists for (some) pronouns, where it indicates a purely possessive relation (which is why it might better be called the possessive case). It does, however, have three different mechanisms for expressing genitive relations in general - the periphrastic genitive, the analytic/agglutinative genitive (both are my own terms, so don't expect to find them in a dictionary), and the possessive.
The periphrastic genitive renders the genitive phrase as a prepositional phrase with the preposition 'of', e.g "fog of war". This is the "true" genitive in English, in that it is the mechanism that can express nearly every possible genitive phrase (although some phrases may sound better using one of the other mechanisms), while the others are much more restricted in use.
The analytic/agglutinative genitive renders the genitive relation purely by shoving two (or more) nouns together, e.g. "milk carton" (compare to "carton of milk"), "wood chips" ("chips of wood"), etc. I call it the analytic genitive because the genitive relation is expressed purely analytically - through word order, rather than word form. I call it the agglutinative genitive because in more synthetic relatives of English, such as German or Old English, the modifying word(s) would be agglutinated with the modified word to form a single large word (this can sometimes produce very long compound words); for example, "girl next door" would be "Nachbarmädchen" in German (literally "neighborgirl"). In English, this type of genitive is restricted to certain types of relations, and is especially used for classification.
Finally, English has a special means of expressing possessive relations, a subset of genitive relations (it's unclear exactly how this form originated; some argue that it's an evolution of the genitive case, while others are more skeptical of that conclusion). Note here that English has a rather broad concept of possession, and as such the possessive can be used with some relations that aren't strictly possessive in nature.
The Genitive in Latin
As a relative of English, the genitive in Latin is much the same, in that genitive relations can be expressed via prepositional phrases (which are substantially similar to their English equivalents) or analytically; however, Latin also has an actual genitive case (a result of which is that it does not need a possessive structure as English does), which is used very commonly. A few examples of the genitive case in Latin:
Finally, As Japanese does not have true case, like English it uses non-case-based structures for constructing the genitive. Japanese has two genitive-marking particles - 'no' and 'na', which differ only in whether the genitive phrase is abstract (e.g. 'baka' - 'stupidity') or concrete (e.g. 'shakunetsu' - 'scorching heat') - used in the periphrastic genitive, as well as the agglutinative/analytic genitive (the genitive may or may not actually be agglutinated). However, Japanese also shows off several uses of the genitive that English and Latin do not.
Consider the rough definition of the genitive - something which specifies or defines a noun phrase; where have we heard this definition before? Well, for one, it's the definition of adjectives. In fact, a language can get by just fine with almost no true adjectives (e.g. the kind we have in English), instead opting to use one or more of the three alternate methods I talked about previously. One of these methods is to use the genitive with descriptive nouns, a method that is used extensively in Japanese and Caia (in fact, it's the primary method in Caia). A couple example from previously used words would be "baka na inu": "stupid dog" (literally "dog of/with stupidity") and"shakunetsu no koi": "scorching love" (literally "love of scorching heat").
The other structure this definition matches is apposition; that is, a restatement of something previously said for the purpose of definition or clarification, such as the italicized part in "Julius, son of Ambrose". In a previous post on the predicative, I explained that in Indo-European languages such phrases are in the same case as the phrase they restate, but that we could imagine other reasonable ways of expressing them. In Japanese such phrases may be expressed either by the analytic genitive (e.g. "kemono [beast] no sousha [player, generally of an instrument] Erin": "Erin the beast player") or the periphrastic genitive (e.g. "Hamelin no violin hiki [player]": "Hamelin the violinist" - literally "violinist of Hamelin").
Finally, some languages, such as Japanese and Caia, allow other relational phrases to be put into the genitive, where the genitive particle indicates the modification of a noun phrase, and the other relational particle specifies the precise relation of the genitive phrase (you might say the genitive particle applies top-down, while the other relational particle applies bottom-up). Some examples of this in Japanese:
Genitive Relation
The genitive relation consists of a noun phrase (the minimum noun phrase in most languages being a single noun) that modifies another noun phrase. As described by grammar books, a genitive phrase is a noun phrase that specifies or describes another noun phrase, with the result being that the genitive phrase helps to answer the question "which?" with regard to the modified noun phrase; e.g. "the girl next door's dog" (genitive in italic) answers "which dog?", "music of Starcraft" answers "which music?", etc.
As it essentially contains all manner of modifying phrases, clearly the genitive is an extremely broad relation, and there are quite a few more-specific relations that fall under the umbrella of "genitive". A semi-exhaustive listing of the various types of relations the genitive can express:
- Physical possession: "my computer"
- Abstract possession: "your happiness"
- Relational: "his wife"
- Compositional: "bottle of water", "box of nails"
- Quantitative: "half of them"
- Subjective: "her snoring"
- Objective: "its destruction"
- Purpose: "sledgehammer of castration"
- Location: "citizens of America"
- Origin: "Kazuhiro Sasaki of Japan" (for those who don't know, he's a Japanese baseball player that was recruited to play professionally in America)
- Affiliation: "Apple's Steve Jobs"
- Attributive: "thing of beauty"
- Topical: "Of Mice and Men"
- Appositional/classifying: "President Obama"
One very good reason for a language to have more than a single mechanism of representing the genitive is that any single such mechanism would necessarily be ambiguous. Take, for example, the English phrase "betrayal of Illidan"; with this phrase it is unclear whether the genitive is subjective or objective - that is, whether Illidan is the betrayer or the one betrayed. Other such examples can be invented based on other uses of the genitive.
Of course it's not all bad; if there were no ambiguity in language, puns would be impossible. One real-life example of genitive ambiguity comes from IRC. jfroy says 'Snow Leopard meeting' in the sense of 'meeting about [Apple] Snow Leopard', to which I reply, interpreting that as 'meeting with a snow leopard':
<@jfroy> Snow Leopard meeting over. Ow.
<enma_hinobara>Next time don't poke it
<enma_hinobara>It's not Kaity
The Genitive in English
It's arguable whether English still has a genitive case; if it does, it only exists for (some) pronouns, where it indicates a purely possessive relation (which is why it might better be called the possessive case). It does, however, have three different mechanisms for expressing genitive relations in general - the periphrastic genitive, the analytic/agglutinative genitive (both are my own terms, so don't expect to find them in a dictionary), and the possessive.
The periphrastic genitive renders the genitive phrase as a prepositional phrase with the preposition 'of', e.g "fog of war". This is the "true" genitive in English, in that it is the mechanism that can express nearly every possible genitive phrase (although some phrases may sound better using one of the other mechanisms), while the others are much more restricted in use.
The analytic/agglutinative genitive renders the genitive relation purely by shoving two (or more) nouns together, e.g. "milk carton" (compare to "carton of milk"), "wood chips" ("chips of wood"), etc. I call it the analytic genitive because the genitive relation is expressed purely analytically - through word order, rather than word form. I call it the agglutinative genitive because in more synthetic relatives of English, such as German or Old English, the modifying word(s) would be agglutinated with the modified word to form a single large word (this can sometimes produce very long compound words); for example, "girl next door" would be "Nachbarmädchen" in German (literally "neighborgirl"). In English, this type of genitive is restricted to certain types of relations, and is especially used for classification.
Finally, English has a special means of expressing possessive relations, a subset of genitive relations (it's unclear exactly how this form originated; some argue that it's an evolution of the genitive case, while others are more skeptical of that conclusion). Note here that English has a rather broad concept of possession, and as such the possessive can be used with some relations that aren't strictly possessive in nature.
The Genitive in Latin
As a relative of English, the genitive in Latin is much the same, in that genitive relations can be expressed via prepositional phrases (which are substantially similar to their English equivalents) or analytically; however, Latin also has an actual genitive case (a result of which is that it does not need a possessive structure as English does), which is used very commonly. A few examples of the genitive case in Latin:
- "agricolae [farmer, genitive] filia [daughter, nominative]": "farmer's daughter"
- "horum [these, genitive] omnium [all, genitive]": "of all these/of all of these"
- "vir [man, nominative] magnae [great, genitive] virtutis [courage, genitive]": "man of great courage"
- "fossa [ditch, nominative] decem [10] pedum [foot, genitive]": "ditch of ten feet/ten-foot ditch"
- "Rex [king, nominative] belli [war, genitive] cupidus [desirous, nominative] est [is]": "The king is desirous of war"
Finally, As Japanese does not have true case, like English it uses non-case-based structures for constructing the genitive. Japanese has two genitive-marking particles - 'no' and 'na', which differ only in whether the genitive phrase is abstract (e.g. 'baka' - 'stupidity') or concrete (e.g. 'shakunetsu' - 'scorching heat') - used in the periphrastic genitive, as well as the agglutinative/analytic genitive (the genitive may or may not actually be agglutinated). However, Japanese also shows off several uses of the genitive that English and Latin do not.
Consider the rough definition of the genitive - something which specifies or defines a noun phrase; where have we heard this definition before? Well, for one, it's the definition of adjectives. In fact, a language can get by just fine with almost no true adjectives (e.g. the kind we have in English), instead opting to use one or more of the three alternate methods I talked about previously. One of these methods is to use the genitive with descriptive nouns, a method that is used extensively in Japanese and Caia (in fact, it's the primary method in Caia). A couple example from previously used words would be "baka na inu": "stupid dog" (literally "dog of/with stupidity") and"shakunetsu no koi": "scorching love" (literally "love of scorching heat").
The other structure this definition matches is apposition; that is, a restatement of something previously said for the purpose of definition or clarification, such as the italicized part in "Julius, son of Ambrose". In a previous post on the predicative, I explained that in Indo-European languages such phrases are in the same case as the phrase they restate, but that we could imagine other reasonable ways of expressing them. In Japanese such phrases may be expressed either by the analytic genitive (e.g. "kemono [beast] no sousha [player, generally of an instrument] Erin": "Erin the beast player") or the periphrastic genitive (e.g. "Hamelin no violin hiki [player]": "Hamelin the violinist" - literally "violinist of Hamelin").
Finally, some languages, such as Japanese and Caia, allow other relational phrases to be put into the genitive, where the genitive particle indicates the modification of a noun phrase, and the other relational particle specifies the precise relation of the genitive phrase (you might say the genitive particle applies top-down, while the other relational particle applies bottom-up). Some examples of this in Japanese:
- "ano [that] group to [with] no kankei [relationship]": "relationship with that group" (literally "relationship of with that group")
- "ano kaisha [company] to no keiyaku [contract]": "contract with that company"
- "Asia kara [from] no ryuugakusei [exchange student]": "exchange student from Asia" (literally "exchange student of from Asia")
- "Hokkaido kara no omiyage [souvenir]": "souvenir from Hokkaido"
- "America ye [towards/to] no monkowohiraku [opening the door; this is rendered as a noun in this phrase, not a verb, in this example]": "opening the door to America" (literally "opening the door of to America")
- "anata [you] ye no tegami [letter]": "letter to/for you"
Wednesday, May 27, 2009
The Dative
The dative, or indirect object, is one of the four core cases - those cases that are almost universal in languages that have case - along with nominative/absolutive, accusative/ergative, and the genitive.
While the details vary by language, the general concept for the dative is a goal or direction of an action (especially with regard to transfer or movement), or the one perceiving an action. In the study of grammatical role, the dative roughly corresponds to the role of goal (although it may have additional uses in a given language). While English has pretty much lost its case system, the logical dative remains, usually expressed with the prepositions 'to' or 'for'.
Some examples of the various uses in several languages from Wikipedia and other sources (note that I'm mainly covering the most common uses, not ones that are specific to certain languages):
Goal of Transfer:
Latin: "Regina puellae pecuniam dat" ("The queen gives money to the girl")
Japanese: "Kodomo ni yaru" ("Give to the child")
Goal of Movement:
Japanese: "Ano hito wa gakkou ni haitta" ("That man has gone to the school")
Goal of Intent or Benefit
Latin: "Auxilio vocare" ("I call for help")
Latin: "Puellae ornamento est" ("It is for the girl's decoration")
Latin: Graecis agros colere ("To till fields for the Greeks")
Greek: "τῷ βασιλεῖ μάχομαι" ("I fight for the king")
Greek: "πᾶς ἀνὴρ αὑτῷ πονεῖ" ("Every man toils for himself")
Goal of Experience:
Latin: "Vir bonus mihi videtur' 'the man seems good to me"
Latin: "Quid mihi Celsus agit?" ("What is Celsus doing [that I am interested in]?")
One of the more peculiar uses (at least for English speakers, as I don't think English has anything like it), is the dative of possession. This renders phrases of possession as phrases of existence with the possessor in the dative. The logical basis of this usage is that in terms of grammatical role, in phrases of possession the possessor is technically classified as the goal.
Some examples:
Latin: "Angelis alae sunt" (literally "For angels there are wings"; freely "Angels have wings")
Greek: "ἄλλοις μὲν γὰρ χρήματα ἐστι πολλὰ καὶ ἵπποι, ἡμῖν δὲ ξύμμαχοι ἀγαθοί" (literally "For others there is a lot of money and ships and horses, but for us there are good allies")
Japanese: "Watashi ni tsuno ga nai" (literally "For me there aren't horns")
Tsez: "Кидбехъор кIетIу зовси" (literally "For the girl there was a cat")
While the details vary by language, the general concept for the dative is a goal or direction of an action (especially with regard to transfer or movement), or the one perceiving an action. In the study of grammatical role, the dative roughly corresponds to the role of goal (although it may have additional uses in a given language). While English has pretty much lost its case system, the logical dative remains, usually expressed with the prepositions 'to' or 'for'.
Some examples of the various uses in several languages from Wikipedia and other sources (note that I'm mainly covering the most common uses, not ones that are specific to certain languages):
Goal of Transfer:
Latin: "Regina puellae pecuniam dat" ("The queen gives money to the girl")
Japanese: "Kodomo ni yaru" ("Give to the child")
Goal of Movement:
Japanese: "Ano hito wa gakkou ni haitta" ("That man has gone to the school")
Goal of Intent or Benefit
Latin: "Auxilio vocare" ("I call for help")
Latin: "Puellae ornamento est" ("It is for the girl's decoration")
Latin: Graecis agros colere ("To till fields for the Greeks")
Greek: "τῷ βασιλεῖ μάχομαι" ("I fight for the king")
Greek: "πᾶς ἀνὴρ αὑτῷ πονεῖ" ("Every man toils for himself")
Goal of Experience:
Latin: "Vir bonus mihi videtur' 'the man seems good to me"
Latin: "Quid mihi Celsus agit?" ("What is Celsus doing [that I am interested in]?")
One of the more peculiar uses (at least for English speakers, as I don't think English has anything like it), is the dative of possession. This renders phrases of possession as phrases of existence with the possessor in the dative. The logical basis of this usage is that in terms of grammatical role, in phrases of possession the possessor is technically classified as the goal.
Some examples:
Latin: "Angelis alae sunt" (literally "For angels there are wings"; freely "Angels have wings")
Greek: "ἄλλοις μὲν γὰρ χρήματα ἐστι πολλὰ καὶ ἵπποι, ἡμῖν δὲ ξύμμαχοι ἀγαθοί" (literally "For others there is a lot of money and ships and horses, but for us there are good allies")
Japanese: "Watashi ni tsuno ga nai" (literally "For me there aren't horns")
Tsez: "Кидбехъор кIетIу зовси" (literally "For the girl there was a cat")
Friday, May 15, 2009
Basic Word Order
Obviously, English typically has has the word order subject-verb[-(direct) object] (or SVO, for short). This can be altered in specific situations due to WH-movement (e.g. OSV in "What would you like?") and other structures, but SVO is the normal word order.
However, this isn't the only possible word order. Basic combination math tells us that there are six possible orders: SVO, SOV (e.g. Japanese, Korean, Latin, and Proto-Indo-European), VSO (e.g. Hebrew, Trique, and Caia), VOS, OSV, OVS; these are, however, not all equally likely.
Two rules have been developed which govern the prominence of different word orders. First, there's a tendency for the subject to precede the (direct) object. This is known as subject salience. There have been various theories on the exact reason for this; the general idea is that it's more natural for the subject to precede the object because they subject is typically the source of an action, and thus precedes the object in both cause and effect and chronological order.
The other is that there is a tendency for the object to sit next to the verb (on either side). Linguistics thus far has developed the notion that the object and the verb logically form a structure called the predicate, which stands apart from the subject (i.e. a clause is typically defined as a subject + a predicate). I haven't investigated the full depth of why this has been decided, so I couldn't really give more detail than that (though it's noteworthy that this idea is consistent with my hypothesis that language began as commands - the command forms the predicate, and the subject was added in later).
Thus, we have four classes of word order: those that meet both conditions, those that meet one or the other, and those that meet neither. In agreement with theory, SVO and SOV are by far the most common among languages, at 42% and 45%, respectively. On the distant second tier is VSO, which places the subject between verb and object, occurring with 9% frequency. On the again distant third tier are VOS, and OVS, each placing the object before the subject; these occur with 3% and 1% frequency, respectively. At the bottom is OSV, which violates both rules, and was not seen in any language in this survey of 402 languages.
Clearly the subject salience rule dominates in significance, as orders where object precedes subject are very rare (3% or less); although it's also true that languages where the object is not next to the verb are uncommon (9% or less).
However, this isn't the only possible word order. Basic combination math tells us that there are six possible orders: SVO, SOV (e.g. Japanese, Korean, Latin, and Proto-Indo-European), VSO (e.g. Hebrew, Trique, and Caia), VOS, OSV, OVS; these are, however, not all equally likely.
Two rules have been developed which govern the prominence of different word orders. First, there's a tendency for the subject to precede the (direct) object. This is known as subject salience. There have been various theories on the exact reason for this; the general idea is that it's more natural for the subject to precede the object because they subject is typically the source of an action, and thus precedes the object in both cause and effect and chronological order.
The other is that there is a tendency for the object to sit next to the verb (on either side). Linguistics thus far has developed the notion that the object and the verb logically form a structure called the predicate, which stands apart from the subject (i.e. a clause is typically defined as a subject + a predicate). I haven't investigated the full depth of why this has been decided, so I couldn't really give more detail than that (though it's noteworthy that this idea is consistent with my hypothesis that language began as commands - the command forms the predicate, and the subject was added in later).
Thus, we have four classes of word order: those that meet both conditions, those that meet one or the other, and those that meet neither. In agreement with theory, SVO and SOV are by far the most common among languages, at 42% and 45%, respectively. On the distant second tier is VSO, which places the subject between verb and object, occurring with 9% frequency. On the again distant third tier are VOS, and OVS, each placing the object before the subject; these occur with 3% and 1% frequency, respectively. At the bottom is OSV, which violates both rules, and was not seen in any language in this survey of 402 languages.
Clearly the subject salience rule dominates in significance, as orders where object precedes subject are very rare (3% or less); although it's also true that languages where the object is not next to the verb are uncommon (9% or less).
Sunday, April 26, 2009
Empathy Hierarchy and Other Ungodly Things
A ways back I wrote a post about empathy (also called animacy or agency) hierarchies and how the empathy hierarchy in Spanish and other romance languages worked, explaining a peculiarity of their grammar that had haunted me since high school Spanish. Unfortunately, I was excited at the time I was writing, and didn't explain it near as clearly as I should have.
Anyway, to briefly recap, empathy hierarchy is a system of classifying entities (e.g. nouns) according to the level of empathy the speaker feels towards them. While the number and composition of levels vary by language, the general trends are 1st person >= 2nd person >= 3rd person, animate >= inanimate, definite >= indefinite (e.g. "those people" >= "some people"). In Spanish there were three levels with regard to verb agreement: 1st person ("[yo] juego" - "I play") > 2nd person familiar ("[tú] juegas") > 2nd person formal ("usted juega") + 3rd person ("[él] juega"); this makes logical sense because you would have more empathy towards those you would address with the 2nd person familiar (e.g. your friends and family) than those you would address with the formal 2nd person (e.g. people you meet on the street). One of the most common things governed by empathy hierarchy is verb agreement, which is exactly what is seen in Spanish in the above examples.
One of the languages (actually a family of languages) I'm creating for one of my stories turns the concept of empathy hierarchy completely upside-down. In these languages, verbs agree with the subject, direct object [if any], and indirect object [if any] (this is as complicated as it sounds, though some real "polypersonal" languages actually do have verbs that agree with all three of those). As the subject, direct object, and indirect object are not marked for case and word order is free, verb agreement is really the only way to tell who's doing what in a sentence; for example, you could have something like "Dog bone boy gaveheitit" or "Bone boy dog gaveheitit", which would mean more or less the same thing, as word order isn't important in this respect (word order indicates more subtle things, such as what is emphasized in a sentence).
However, instead of relying on animacy or empathy to create this agreement system (animacy and/or empathy being pretty much universal in natural languages, as far as I know), these languages use rank. Specifically, a seven-tiered hierarchy with levels I label -3 (lowest) to +3 (highest). For humans, rank reflects social status relative to the speaker. Rank 0 is reserved for the first person, ranks below 0 represent those below the social status of the speaker, and ranks above 0 above the speaker, with the degree of difference indicated by the rank (rank x + 1 > rank x); e.g. use of rank 2 by the speaker to refer to someone would mean that that person has a social status substantially higher than that of the speaker. For nonhumans (animals and inanimate objects) rank is based on respect for that thing - positive and/or honorable things would have high rank, negative or dishonorable things would have low rank; in this way rank resembles more arbitrary noun class systems.
Now, depending on how deeply you thought about what I just said, you may or may not already be scared of this idea; so let me illustrate. Suppose you have two coworkers at a job talking to each other (thus they'd be of the same social status). As it's customary to show some respect to those of the same rank as you in East Asian cultures (these languages are loosely based on Japanese, with a bunch of fun stuff thrown in), both would refer to each other ("you" in English) with the +1 level. However, as this system does not distinguish between 2nd and 3rd person, if one of them uses the +1 level for something, it could just as easily refer to some third co-worker who wasn't a part of the conversation (a 3rd person, no pun intended). Or it could simply refer to some inanimate object that just happens to be at the +1 rank.
But that's easy stuff. Now imagine a person writing to a politician. If this was just an average person, they might use rank +2 to refer to the politician ("you", in English), given the significant difference in social status. The politician would then use rank -2 to refer to the person (or -1, if the politician wanted to be polite), which would also be written "you" in English. Now suppose they were talking about the king. The lay person would probably refer to the king with +3 rank; the politician, however, would use a lower rank (+1 or +2), as their own social status is greater. Of course, any of these (-2, -1, +1, +2, +3) could simply be referring to an inanimate object, instead.
Worse yet, the innate rank of an animal or inanimate object could be modified based on the rank of a possessor; that is, if the possessor of an object has a significantly different rank than the speaker, the rank of the thing possessed might be elevated or lowered to reflect the possessor.
Suppose that dog' has an innate rank of +1. In the previous scenario, the person might then use rank +1 to refer to his dog, +2 to refer to the politician's dog, and +2 or +3 to refer to the king's dog. The politician, on the other hand, might refer to his own dog as +1, the person's dog as -1, and the king's dog as +2.
And as one last nail in the coffin of sanity is a concept I call the heterogeneous plural. That is, a plural that comprises members of significantly different rank. In this case, the rank of the group as a whole would be the rank of the highest-ranking member. So if the person writing a politician referred to level +3, he could be referring the the king, the king and the politician together, the king's dog, or something else entirely.
That's me: raping your mind since 1983.
Anyway, to briefly recap, empathy hierarchy is a system of classifying entities (e.g. nouns) according to the level of empathy the speaker feels towards them. While the number and composition of levels vary by language, the general trends are 1st person >= 2nd person >= 3rd person, animate >= inanimate, definite >= indefinite (e.g. "those people" >= "some people"). In Spanish there were three levels with regard to verb agreement: 1st person ("[yo] juego" - "I play") > 2nd person familiar ("[tú] juegas") > 2nd person formal ("usted juega") + 3rd person ("[él] juega"); this makes logical sense because you would have more empathy towards those you would address with the 2nd person familiar (e.g. your friends and family) than those you would address with the formal 2nd person (e.g. people you meet on the street). One of the most common things governed by empathy hierarchy is verb agreement, which is exactly what is seen in Spanish in the above examples.
One of the languages (actually a family of languages) I'm creating for one of my stories turns the concept of empathy hierarchy completely upside-down. In these languages, verbs agree with the subject, direct object [if any], and indirect object [if any] (this is as complicated as it sounds, though some real "polypersonal" languages actually do have verbs that agree with all three of those). As the subject, direct object, and indirect object are not marked for case and word order is free, verb agreement is really the only way to tell who's doing what in a sentence; for example, you could have something like "Dog bone boy gaveheitit" or "Bone boy dog gaveheitit", which would mean more or less the same thing, as word order isn't important in this respect (word order indicates more subtle things, such as what is emphasized in a sentence).
However, instead of relying on animacy or empathy to create this agreement system (animacy and/or empathy being pretty much universal in natural languages, as far as I know), these languages use rank. Specifically, a seven-tiered hierarchy with levels I label -3 (lowest) to +3 (highest). For humans, rank reflects social status relative to the speaker. Rank 0 is reserved for the first person, ranks below 0 represent those below the social status of the speaker, and ranks above 0 above the speaker, with the degree of difference indicated by the rank (rank x + 1 > rank x); e.g. use of rank 2 by the speaker to refer to someone would mean that that person has a social status substantially higher than that of the speaker. For nonhumans (animals and inanimate objects) rank is based on respect for that thing - positive and/or honorable things would have high rank, negative or dishonorable things would have low rank; in this way rank resembles more arbitrary noun class systems.
Now, depending on how deeply you thought about what I just said, you may or may not already be scared of this idea; so let me illustrate. Suppose you have two coworkers at a job talking to each other (thus they'd be of the same social status). As it's customary to show some respect to those of the same rank as you in East Asian cultures (these languages are loosely based on Japanese, with a bunch of fun stuff thrown in), both would refer to each other ("you" in English) with the +1 level. However, as this system does not distinguish between 2nd and 3rd person, if one of them uses the +1 level for something, it could just as easily refer to some third co-worker who wasn't a part of the conversation (a 3rd person, no pun intended). Or it could simply refer to some inanimate object that just happens to be at the +1 rank.
But that's easy stuff. Now imagine a person writing to a politician. If this was just an average person, they might use rank +2 to refer to the politician ("you", in English), given the significant difference in social status. The politician would then use rank -2 to refer to the person (or -1, if the politician wanted to be polite), which would also be written "you" in English. Now suppose they were talking about the king. The lay person would probably refer to the king with +3 rank; the politician, however, would use a lower rank (+1 or +2), as their own social status is greater. Of course, any of these (-2, -1, +1, +2, +3) could simply be referring to an inanimate object, instead.
Worse yet, the innate rank of an animal or inanimate object could be modified based on the rank of a possessor; that is, if the possessor of an object has a significantly different rank than the speaker, the rank of the thing possessed might be elevated or lowered to reflect the possessor.
Suppose that dog' has an innate rank of +1. In the previous scenario, the person might then use rank +1 to refer to his dog, +2 to refer to the politician's dog, and +2 or +3 to refer to the king's dog. The politician, on the other hand, might refer to his own dog as +1, the person's dog as -1, and the king's dog as +2.
And as one last nail in the coffin of sanity is a concept I call the heterogeneous plural. That is, a plural that comprises members of significantly different rank. In this case, the rank of the group as a whole would be the rank of the highest-ranking member. So if the person writing a politician referred to level +3, he could be referring the the king, the king and the politician together, the king's dog, or something else entirely.
That's me: raping your mind since 1983.
Friday, April 10, 2009
Random Late-Night Thought
For a while I've been aware of a particular piece of linguistic evidence - namely, that cross-linguistically it is common for the imperative (command) form of verbs to be shorter than other forms, seemingly lacking inflectional affixes applied to other conjugations. For example, in Old English, the verb 'creopan' (infinitive form) is 'creope' for first person singular, 'criepth' for third person singular, but 'creop' for imperative singular. This led me to hypothesize that language might have begun as commands, and later evolved to support more general types of expressions by the addition of affixes or extra words.
Think about the significance of this for a moment. One of the most frequent differences between normal sentences and imperative sentences is that the imperative usually does not have a (stated) subject, while, depending on the language, general sentences may require subjects.
Now, one of the big mysteries of linguistics is how we came to have such radically different language systems as accusative, ergative, and topic-comment*. Yet if commands were all that was originally spoken, this provides us with a trivial answer: initially, there was only a direct object and no subject, thus that language would predate the differentiation of the three types.
As the language evolved further, eventually there would be the need to add in a subject; how exactly this was handled would then determine which of the three paths was taken. Accusative languages would place the subject in a separate case (nominative) from the direct object (accusative case). Ergative languages would classify the subject based on whether its role is the agent (ergative case) or patient (absolutive case). Finally, topic-comment languages would place the subject (the topic) completely apart from the rest of the sentence (the comment).
*Since I don't think I've talked too much about topic-comment structure, I'll briefly explain here. In topic-comment languages, a topic is stated for a sentence or set of sentences, then a number of comments are made regarding that topic. Japanese, Korean, and Chinese are like this, among others, although it's also possible to use a periphrastic form in languages like English (e.g. "As for the movie [the topic], we'll meet at 2 [the comment]").
Of particular relevance, one thing the topic can be used for is the subject of the sentence, e.g. "As for him, he'll be coming later" (though true topic-comment languages usually wouldn't duplicate the subject as English does - it would be more like "As for him, will come later"); this is frequently done in Japanese and Korean, for example. Of course, the comment may have a different subject than the topic, so topic-comment languages may also be accusative or ergative (e.g. Japanese is accusative). Here I am hypothesizing that initially the subject was represented exclusively as the topic, then further evolution allowed the subject to be within the comment itself (although whether this is true is relatively unimportant to the theory that languages began as commands).
Think about the significance of this for a moment. One of the most frequent differences between normal sentences and imperative sentences is that the imperative usually does not have a (stated) subject, while, depending on the language, general sentences may require subjects.
Now, one of the big mysteries of linguistics is how we came to have such radically different language systems as accusative, ergative, and topic-comment*. Yet if commands were all that was originally spoken, this provides us with a trivial answer: initially, there was only a direct object and no subject, thus that language would predate the differentiation of the three types.
As the language evolved further, eventually there would be the need to add in a subject; how exactly this was handled would then determine which of the three paths was taken. Accusative languages would place the subject in a separate case (nominative) from the direct object (accusative case). Ergative languages would classify the subject based on whether its role is the agent (ergative case) or patient (absolutive case). Finally, topic-comment languages would place the subject (the topic) completely apart from the rest of the sentence (the comment).
*Since I don't think I've talked too much about topic-comment structure, I'll briefly explain here. In topic-comment languages, a topic is stated for a sentence or set of sentences, then a number of comments are made regarding that topic. Japanese, Korean, and Chinese are like this, among others, although it's also possible to use a periphrastic form in languages like English (e.g. "As for the movie [the topic], we'll meet at 2 [the comment]").
Of particular relevance, one thing the topic can be used for is the subject of the sentence, e.g. "As for him, he'll be coming later" (though true topic-comment languages usually wouldn't duplicate the subject as English does - it would be more like "As for him, will come later"); this is frequently done in Japanese and Korean, for example. Of course, the comment may have a different subject than the topic, so topic-comment languages may also be accusative or ergative (e.g. Japanese is accusative). Here I am hypothesizing that initially the subject was represented exclusively as the topic, then further evolution allowed the subject to be within the comment itself (although whether this is true is relatively unimportant to the theory that languages began as commands).
Thursday, March 19, 2009
Random Linguistic Fact of the Day
Technically English (and I think all Germanic languages) doesn't have a future tense. The future is rendered as a mood in English, using a modal (mood auxiliary) verb, in the same class as (and mutually exclusive with) "can", "may", "must", "would", etc. The past and present tenses, on the other hand, are true tenses, and both in the indicative mood (no modal verb or the "do" dummy modal verb).
The logical basis for this distinction has to do with the concept of realis. Essentially that means what it looks like: realis moods have to do with 'real' things - things which are considered certain to have already happened; while irrealis moods are not certain for one reason or another. There's a general tendency in language to regard the future as inherantly uncertain, and thus place it in an irrealis mood.
Whether this is a peculiarity of Germanic languages or is universal among Indo-European languages is unclear. In Latin there is a future tense for the imperative mood (commands - an irealis mood) as well as indicative (events that are certain - a realis mood), but not for subjunctive or supine, two other irealis moods.
For trivia value: Caia does not have tense; aspect and mood are used to imply tense, and if tense must be made absolutely certain, it can be indicated with adverbs. It has three basic moods (more complex moods are specified with helper verbs or particles): indicative, potential, and hypothetical. As in English, the indicative is used for events considered certain, and is used primarily for past and present tense. Potential mood indicates that an event is possible, but not certain; it is used for the future, among other things (although the preferred method of referring to the future is to reduce it to a certain, indicative present expression such as "I intend to go" or "I want to go", which is more precise). The hypothetical refers to events that are known to be false (hence talking about a hypothetical, counter-factual "what if" situation).
The logical basis for this distinction has to do with the concept of realis. Essentially that means what it looks like: realis moods have to do with 'real' things - things which are considered certain to have already happened; while irrealis moods are not certain for one reason or another. There's a general tendency in language to regard the future as inherantly uncertain, and thus place it in an irrealis mood.
Whether this is a peculiarity of Germanic languages or is universal among Indo-European languages is unclear. In Latin there is a future tense for the imperative mood (commands - an irealis mood) as well as indicative (events that are certain - a realis mood), but not for subjunctive or supine, two other irealis moods.
For trivia value: Caia does not have tense; aspect and mood are used to imply tense, and if tense must be made absolutely certain, it can be indicated with adverbs. It has three basic moods (more complex moods are specified with helper verbs or particles): indicative, potential, and hypothetical. As in English, the indicative is used for events considered certain, and is used primarily for past and present tense. Potential mood indicates that an event is possible, but not certain; it is used for the future, among other things (although the preferred method of referring to the future is to reduce it to a certain, indicative present expression such as "I intend to go" or "I want to go", which is more precise). The hypothetical refers to events that are known to be false (hence talking about a hypothetical, counter-factual "what if" situation).
Sunday, March 08, 2009
Random Linguistic Fact of the Day
Ever wonder why it's fairly common to create compound nouns from phrases (e.g. bird-watching, card-carrying), but in all these cases the object comes before the verb participle? Based on English word order it should be watching-bird, etc., yet it never is.
This is probably due to the fact that word order in Proto-Indo-European was very different than the word order used in English and most other Indo-European languages today. In particular, instead of the subject-verb-object order typically used today, PIE (along with more recent ones, like Latin) preferred the subject-object-verb word order. So you might say things like "Avem [bird] spectabam [I watched]" in Latin, which is exactly the order seen in the compounds.
This is probably due to the fact that word order in Proto-Indo-European was very different than the word order used in English and most other Indo-European languages today. In particular, instead of the subject-verb-object order typically used today, PIE (along with more recent ones, like Latin) preferred the subject-object-verb word order. So you might say things like "Avem [bird] spectabam [I watched]" in Latin, which is exactly the order seen in the compounds.
Wednesday, July 09, 2008
& Fun with Turkish
Just saw this amusing segment in the Wikipedia page for Turkish grammar:
Avrupa (Europe)
Avrupalı (European)
Avrupalılaş (become European)
Avrupalılaştır (Europeanize)
Avrupalılaştırama (cannot Europeanize)
Avrupalılaştıramadık (whom [someone] could not Europeanize)
Avrupalılaştıramadıklar (those whom [someone] could not Europeanize)
Avrupalılaştıramadıklarımız (those whom we could not Europeanize)
Avrupalılaştıramadıklarımızdan (one of those whom we could not Europeanize)
Avrupalılaştıramadıklarımızdan mı? (one of those whom we could not Europeanize?)
Avrupalılaştıramadıklarımızdan mısınız? (Are you one of those whom we could not Europeanize?)
You now know the meaning of 'highly agglutinative language'.
Avrupa (Europe)
Avrupalı (European)
Avrupalılaş (become European)
Avrupalılaştır (Europeanize)
Avrupalılaştırama (cannot Europeanize)
Avrupalılaştıramadık (whom [someone] could not Europeanize)
Avrupalılaştıramadıklar (those whom [someone] could not Europeanize)
Avrupalılaştıramadıklarımız (those whom we could not Europeanize)
Avrupalılaştıramadıklarımızdan (one of those whom we could not Europeanize)
Avrupalılaştıramadıklarımızdan mı? (one of those whom we could not Europeanize?)
Avrupalılaştıramadıklarımızdan mısınız? (Are you one of those whom we could not Europeanize?)
You now know the meaning of 'highly agglutinative language'.
Tuesday, June 24, 2008
Random Thought of the Day
Did you ever notice that, in English, the simple past (e.g. "he wrote") and past progressive (e.g. "he was writing") are both very common, yet in the present tense, the present progressive (e.g. "he is writing") is overwhelmingly more common than the simple present (e.g. "he writes")? This fact actually leads into an important linguistic principle, which I'll probably write a post about in the future. I'll just leave it as food for thought, for now.
Monday, June 16, 2008
Cases, Ergative, & Accusative
Something that I vaguely implied previously, but I don't think actually said, was that there is a difference between roles and cases (even worse, there are multiple things that "role" could refer to). Roles are, in theory, purely rational, language-independent categories which describe how nouns relate to their clause's verb. Cases, on the other hand, are language-dependent categories representing many things, and there is rarely (if ever) a 1:1 mapping of the two for a language.
The Grammer of Discourse hypothesizes at least ten universal roles, which I'll only briefly describe.
Experiencer: the person experiencing an emotion or sensation
Patient: the one an action acts on
Agent: the one willfully performing an action
Range: an extension of the verb, such as indicating how, e.g. "Your blood smells good"
Measure: an extension of the verb indicating how much, e.g. "I was only bitten a little bit" (these examples brought to you by Vampire Knight)
Instrument: something which is used to perform an action; this can also be used for animate entities who unintentionally perform an action
Locative: the location an action occurs at
Source: the starting point of some kind of movement or transfer
Goal: the ending point of movement or transfer
Path: the path taken during movement or transfer
If we were to compare this list of roles with typical use of the Latin cases, we would get the following. Note that this list is approximate, and some of the roles like measure and range I'm not even sure how to represent in Latin.
Nominative case: agent, patient, experiencer, instrument
Genitive: unrelated to role in the sentence (roles refer to relation with the verb, not with other nouns)
Dative: goal, patient
Accusative: patient, experiencer, goal, rarely source
Ablative: source, instrument, locative, goal, path, possibly range and measure (some of those requiring prepositions)
Locative (rare): locative
Vocative: not related to role
However, while case is language-specific, some themes (common cases) occur much more often than others. Of the Latin cases, the nominative, genitive, dative, and accusative occur very frequently in all languages; this is not surprising, as these seem the most essential to language in general (though note that they are not guaranteed to mean exactly the same thing in all languages).
The nominative case is roughly defined as the subject of the verb. For transitive verbs having a direct object, the subject is the one performing the action (e.g. "He poked her"); for intransitive verbs the subject is the single argument (e.g. "He was hit"). The accusative case is the object of transitive verbs. Any language having this structure is called a nominative-accusative (or sometimes just accusative) language (which we're going to call N/A in the rest of this post).
However, two others - the ergative and the absolutive - also occur very commonly in languages. The ergative case is defined as the subject of transitive verbs. The absolutive case, however, includes both the subject of intransitive verbs and the object of transitive verbs. Languages using this system are called ergative-absolutive (or sometimes just ergative; E/A, here).
At first this seems very strange and arbitrary - splitting the subject depending on whether the verb is transitive or intransitive. However, this is due to the fact that we don't speak an language. In fact, even the word 'subject' reflects this bias in thinking. The N/A split carries the paradigm that all actions are done by somebody/something, regardless of whether the action is intentional or unintentional, or even whether there's anyone performing the action at all (e.g. in "He fell"). This is called the subject, and for transitive verbs, the one acted on is the called the object; thus the N/A split actually corresponds to a subject/object division.
However, we get a different picture if we discard this assumption and look at things from the perspective of roles. In reality, with many intransitive verbs (such as the one shown above) the "subject" is not the one doing the action at all, but rather the one who is subjected to the action - the patient. Thus the E/A split is based on the paradigm that the ergative case is the doer (agent or instrument) of the action, while the absolutive case is the patient of the action - an agent/patient separation. Taking it one step further, some E/A languages even require that the ergative argument commit the action intentionally, and use a different sentence structure to indicate otherwise (e.g. split-intransitivity languages use either the ergative or absolutive case for the subject of intransitive verbs, depending on whether the action is intentional or not; others use the passive voice for unintentional actions; etc.).
Given this, both seem equally sensible, and the choice itself now seems arbitrary. It's worth noting, also, that most languages in the world are either N/A or E/A. Languages using other systems are rare, which might suggest that the N/A and E/A splits are more sensible and/or useful than other methods. But hold onto that thought.
The Grammer of Discourse hypothesizes at least ten universal roles, which I'll only briefly describe.
Experiencer: the person experiencing an emotion or sensation
Patient: the one an action acts on
Agent: the one willfully performing an action
Range: an extension of the verb, such as indicating how, e.g. "Your blood smells good"
Measure: an extension of the verb indicating how much, e.g. "I was only bitten a little bit" (these examples brought to you by Vampire Knight)
Instrument: something which is used to perform an action; this can also be used for animate entities who unintentionally perform an action
Locative: the location an action occurs at
Source: the starting point of some kind of movement or transfer
Goal: the ending point of movement or transfer
Path: the path taken during movement or transfer
If we were to compare this list of roles with typical use of the Latin cases, we would get the following. Note that this list is approximate, and some of the roles like measure and range I'm not even sure how to represent in Latin.
Nominative case: agent, patient, experiencer, instrument
Genitive: unrelated to role in the sentence (roles refer to relation with the verb, not with other nouns)
Dative: goal, patient
Accusative: patient, experiencer, goal, rarely source
Ablative: source, instrument, locative, goal, path, possibly range and measure (some of those requiring prepositions)
Locative (rare): locative
Vocative: not related to role
However, while case is language-specific, some themes (common cases) occur much more often than others. Of the Latin cases, the nominative, genitive, dative, and accusative occur very frequently in all languages; this is not surprising, as these seem the most essential to language in general (though note that they are not guaranteed to mean exactly the same thing in all languages).
The nominative case is roughly defined as the subject of the verb. For transitive verbs having a direct object, the subject is the one performing the action (e.g. "He poked her"); for intransitive verbs the subject is the single argument (e.g. "He was hit"). The accusative case is the object of transitive verbs. Any language having this structure is called a nominative-accusative (or sometimes just accusative) language (which we're going to call N/A in the rest of this post).
However, two others - the ergative and the absolutive - also occur very commonly in languages. The ergative case is defined as the subject of transitive verbs. The absolutive case, however, includes both the subject of intransitive verbs and the object of transitive verbs. Languages using this system are called ergative-absolutive (or sometimes just ergative; E/A, here).
At first this seems very strange and arbitrary - splitting the subject depending on whether the verb is transitive or intransitive. However, this is due to the fact that we don't speak an language. In fact, even the word 'subject' reflects this bias in thinking. The N/A split carries the paradigm that all actions are done by somebody/something, regardless of whether the action is intentional or unintentional, or even whether there's anyone performing the action at all (e.g. in "He fell"). This is called the subject, and for transitive verbs, the one acted on is the called the object; thus the N/A split actually corresponds to a subject/object division.
However, we get a different picture if we discard this assumption and look at things from the perspective of roles. In reality, with many intransitive verbs (such as the one shown above) the "subject" is not the one doing the action at all, but rather the one who is subjected to the action - the patient. Thus the E/A split is based on the paradigm that the ergative case is the doer (agent or instrument) of the action, while the absolutive case is the patient of the action - an agent/patient separation. Taking it one step further, some E/A languages even require that the ergative argument commit the action intentionally, and use a different sentence structure to indicate otherwise (e.g. split-intransitivity languages use either the ergative or absolutive case for the subject of intransitive verbs, depending on whether the action is intentional or not; others use the passive voice for unintentional actions; etc.).
Given this, both seem equally sensible, and the choice itself now seems arbitrary. It's worth noting, also, that most languages in the world are either N/A or E/A. Languages using other systems are rare, which might suggest that the N/A and E/A splits are more sensible and/or useful than other methods. But hold onto that thought.
Thursday, June 12, 2008
Case & Other Cases
One thing necessary in all languages is that the nouns in a sentence that play various roles/cases must be identifiable. While the exact amount of precision varies by language and by sentence structure (there may be more than one way to say something, or only certain structures may be used in certain cases), all languages have a way to indicate the subject, direct object, etc. (although of course the exact set of roles that exists varies by language, as well). As far as I'm aware, there are three methods of accomplishing this: dependent-marking, head-marking, and analysis (note that none of these terms refers exclusively to role; I'm merely discussing them in this one specific context).
Let's start with the easy one: analysis. This is the method English uses for its core roles: subject, direct object, and sometimes the indirect object. As I pointed out in The Decline of the English Language, Modern English has a fairly rigid word order for its core roles: Subject Verb [IndirectObject] [DirectObject], as in "The boy gave the dog a bone"; some other word orders are used by native speakers, but they're uncommon, and generally only used in certain specific contexts (e.g. the Verb Subject Complement order in "Are you an idiot?"). Thus analysis refers to the use of strict word ordering to determine what role each noun has.
As I mentioned in the same paper, English wasn't always this way: it belongs to the same language family as Latin, all traditionally using dependent-marking of case. Dependent marking refers to the fact that each word is marked to indicate its role. In the same sentence "Puer [boy] cani [dog] os [bone] dabat [gave]", the four words may be placed in any order, and the meaning will still be clear, because the nouns carry the nominative, dative, and accusative cases, respectively (actually, that isn't 100% true; because some cases decline the same way, there can be some ambiguity here).
You might notice that English also does this for non-core roles, which corresponds to greater freedom as to word order. As dependent-marking does not require that the mark actually be attached to the word, English uses prepositions to mark non-core roles, rather than the traditional suffixes of Indo-European languages. This system is used for such roles as instrument in "The boy poked the dog with a bone" (the Latin version, "Puer canes osse pungebat", uses the ablative case, and the accusative case for the dog), the benefactor in "The boy bought a dog for her" (in the Latin version "Puer canes per ea emebat", a preposition is used with the ablative in this case), etc. The last example also illustrates that Latin uses prepositions as well, to mark roles outside the 6 core cases.
Both of those have been something that isn't entirely unfamiliar to English speakers. Even case still (barely) exists in the pronouns and nouns of English (having three and two cases, respectively); the third method, head-marking or agreement, is also not absolutely foreign, though it is uncommon in modern English. Verbs in Indo-European languages traditionally agree with the subject of the sentence - the verbs themselves indicate the grammatical person and number of the subject. While English has all but lost this form of agreement, you can still see vestiges of it. The verb 'am' uniquely identifies the subject as first person singular, while 'is' identifies the subject as third person singular ('are' is ambiguous, because it could refer either to second person singular or any person plural); similarly, the -s form of all other verbs (e.g. 'gives') identify the subject as third person singular. Romance languages like Spanish still contain robust subject-verb agreement, such that it is possible to uniquely identify the subject as first, second, or third person (never mind the bad terminology for now) and singular or plural.
However, you might have noticed something: in languages like Latin that have subject agreement, marking nouns with the nominative case (used for the subject) can be redundant. Head-marking, or polysynthetic, languages do away with this use of case, and purely rely on verb agreement to indicate which nouns have each role. I can't find a good example of a sentence that would indicate how this would work without introducing other things I don't want to get into, so I'm gonna make one up:
Let's start with the easy one: analysis. This is the method English uses for its core roles: subject, direct object, and sometimes the indirect object. As I pointed out in The Decline of the English Language, Modern English has a fairly rigid word order for its core roles: Subject Verb [IndirectObject] [DirectObject], as in "The boy gave the dog a bone"; some other word orders are used by native speakers, but they're uncommon, and generally only used in certain specific contexts (e.g. the Verb Subject Complement order in "Are you an idiot?"). Thus analysis refers to the use of strict word ordering to determine what role each noun has.
As I mentioned in the same paper, English wasn't always this way: it belongs to the same language family as Latin, all traditionally using dependent-marking of case. Dependent marking refers to the fact that each word is marked to indicate its role. In the same sentence "Puer [boy] cani [dog] os [bone] dabat [gave]", the four words may be placed in any order, and the meaning will still be clear, because the nouns carry the nominative, dative, and accusative cases, respectively (actually, that isn't 100% true; because some cases decline the same way, there can be some ambiguity here).
You might notice that English also does this for non-core roles, which corresponds to greater freedom as to word order. As dependent-marking does not require that the mark actually be attached to the word, English uses prepositions to mark non-core roles, rather than the traditional suffixes of Indo-European languages. This system is used for such roles as instrument in "The boy poked the dog with a bone" (the Latin version, "Puer canes osse pungebat", uses the ablative case, and the accusative case for the dog), the benefactor in "The boy bought a dog for her" (in the Latin version "Puer canes per ea emebat", a preposition is used with the ablative in this case), etc. The last example also illustrates that Latin uses prepositions as well, to mark roles outside the 6 core cases.
Both of those have been something that isn't entirely unfamiliar to English speakers. Even case still (barely) exists in the pronouns and nouns of English (having three and two cases, respectively); the third method, head-marking or agreement, is also not absolutely foreign, though it is uncommon in modern English. Verbs in Indo-European languages traditionally agree with the subject of the sentence - the verbs themselves indicate the grammatical person and number of the subject. While English has all but lost this form of agreement, you can still see vestiges of it. The verb 'am' uniquely identifies the subject as first person singular, while 'is' identifies the subject as third person singular ('are' is ambiguous, because it could refer either to second person singular or any person plural); similarly, the -s form of all other verbs (e.g. 'gives') identify the subject as third person singular. Romance languages like Spanish still contain robust subject-verb agreement, such that it is possible to uniquely identify the subject as first, second, or third person (never mind the bad terminology for now) and singular or plural.
However, you might have noticed something: in languages like Latin that have subject agreement, marking nouns with the nominative case (used for the subject) can be redundant. Head-marking, or polysynthetic, languages do away with this use of case, and purely rely on verb agreement to indicate which nouns have each role. I can't find a good example of a sentence that would indicate how this would work without introducing other things I don't want to get into, so I'm gonna make one up:
In this example, theyare attachedit pronouns representingthem the subject, direct object, and indirect object the verb of each clause. As with English pronouns in general, theyagree the attached pronouns with number and gender of the nouns. For the verbs, iusedthem the subject-verb-object order and pronoun cases, to makethem the verbs easier to read for English speakers. However, iusedthem varying word orders for nouns in the clauses to illustrateit how itcan be used head-marking with different word orders. Typically theywould useit head-marking head-marking languages with other modifiers like possessives, as well.Finally, the Totonac language takes polysynthesis to a ridiculous extreme. According to the examples in The Grammar of Discourse, Totonac merely lists all roles in the sentence, without using agreement to indicate which nouns have which roles. One example given (I'm kind of making up my own orthography, here) is "liiteemaktamaahua [literally 'with-passing by-from-buy'] tumin [money]", which means "As [he] passes by, [he] buys [it] from [him] with money". Amazingly (and completely against expectations), native speakers of Totonac can actually understand each other.
Sunday, May 04, 2008
Orthographic Language Identification Using Artificial Neural Networks
Finally finished adding in the figures from the presentation that weren't in the original version of the paper. I have no idea why those three Visio figures are so ugly; they don't look that way in Visio, Word, and Powerpoint, but magically turn ugly when printed with PDF reDirect (which has worked well before those three figures).
Orthographic Language Identification Using Artificial Neural Networks
I'd like to post the full source to TextBreaker, but a lot of it was written in a hurry and needs cleaning up and commenting. Combine my infamous laziness (and past experience with MPQDraft) and time will tell if I ever get around to it :P
Orthographic Language Identification Using Artificial Neural Networks
I'd like to post the full source to TextBreaker, but a lot of it was written in a hurry and needs cleaning up and commenting. Combine my infamous laziness (and past experience with MPQDraft) and time will tell if I ever get around to it :P
Labels:
linguistics,
programming,
reallife
Saturday, May 03, 2008
The Grand Unification
The assumptions you start with can often limit the set of conclusions you are able to arrive at through logical reasoning. Computer science people were having a heck of a time attempting to adapt the binary tree - the standard for in-memory search structures - to media with large seek time, particularly disk drives. Progress was slow and fairly unproductive while the basic definition of binary tree held. Finally, somebody questioned the assumption and thought: what if we built the tree from the bottom up? And so the B-tree was born, and remains the standard in index structures to this day, with incremental improvement.
Other times, when creating, the assumptions themselves are fun to play with and observe the results. In Caia, I initially envisioned all of the major words - nouns, adjectives, and verbs - to be nouns. Nouns became adjectives when used with an attributive particle analogous to "having" (e.g. "woman having beauty" vs. "beautiful woman"). Taking an idea from Japanese, nouns became verbs by an auxiliary verb meaning more or less "do" or "make" (e.g. "make cut" vs. "cut"). Unfortunately, I ultimately concluded that verbs had to be separate, due to both practical concerns (specifically, concerns about making thoughts too long) and theoretical concerns (different nouns require different semantics, which would produce inconsistent theoretical behavior).
With another one of my languages, I took a different route, with some very interesting results. As this language is synthetic (unlike Caia, which is strongly analytic and isolating), I had quite a bit more flexibility. This language was actually modeled on the Altaic languages - Japanese, Korean, Mongolian, Turkish, etc.; as such, I suppose I can't claim that I invented this (what I'm getting to), but merely took what existed in Altaic languages and perfected it to a degree that doesn't exist in nature - at least to my knowledge.
The result is a language is which nouns, adjectives, and verbs are, in fact, all verbs; they are conjugated exactly the same, and play the same grammatical role. Even though this particular language is fictional, and I'm not expecting anybody to actually speak it, this idea might show up in other languages of mine (possibly a real one), as I find it exceptionally elegant. However, I do like to refer to "nouns" as substantives, "adjectives" as attributes, and "verbs" as actions; this is done because there are slightly differences in meaning between the noun bases in the three cases.
I explained previously that Japanese verbs had several base forms - usually distinguished by one vowel - which were then agglutinated with other things to form complex verb forms. I won't describe them again, as I use different names for the ones in my language, which might result in confusion. In my language, there are at least five different base forms of each verb: neutral, conclusive, attributive, participial, instantiative, and conjunctive (note that these names are not final, and I'm open to suggestions).
The conclusive form is the same as the Japanese conclusive: it's the main verb of the last clause in a sentence; it indicates the end of the sentence, and often has a number of suffixes indicating various details about the sentence. The attributive form is the same as one of the two uses of the Japanese attributive: it's the verb of a relative clause; adjectives are attached to nouns by forming relative clauses (e.g. "cat that is fat" vs. "fat cat"). The participial form resembles the second use of the Japanese attributive form: it is a noun referring to the act named by the verb (essentially an English gerund or participle, as in "watching FLCL makes me want to kill people"). The instantiative is unique to my language, and refers to an instance of the verb's action; for example, "a run" would be an instance of the verb "run". The conjunctive is similar to the Japanese -te form (and also includes the Japanese conjunctive base); specifically, it is used in verbs not in the last clause of a sentence (I'll come back to that). Finally, the neutral form is used in agglutination, and has something of a flexible, context-dependent meaning.
So, how exactly does this framework allow the grand unification? Let's look at an example of the specific meanings of the different bases for each word type (substantive, attribute, and action). Although I should note that not every base is necessary to unify the three; some are simply part of the bigger picture for the language.
For the substantive "human":
Conclusive form: "is [a] human"
Attributive form: "who is [a] human"
Participial: "being human"
Instantiative: "human"
Conjunctive: "is [a] human, and..."
For the attribute "fat":
Conclusive form: "is fat"
Attributive form: "who is fat"/"fat"
Participial form: "being fat"
Instantiative form: "fat thing/person"
Conjunctive form: "is fat, and..."
For the action "travel":
Conclusive form: "travel"
Attributive form: "who travels"
Participial form: "traveling"
Instantiative form: "journey"
Conjunctive form: "travel, and..."
Thus we are able to use identical conjugation for each type of word, treating the first two as stative verbs and the last as an active verb, in an elegant unified system. The real key to this, I think, was the separation of participial and instantiative forms. Note that not all actions have an instantiative form; it only exists where it makes sense: where something is produced or performed.
Now that I've explained how the unification works, there's just one more loose end to tie up: the meaning of the conjunctive form. This form is somewhat foreign to English speakers, because English only works this way in one of the circumstances this language uses it for (specifically, conjunction - e.g. "Murasaki is seven years old and [is] surprisingly well-spoken").
In English, when we have multiple clauses in a sentence that are related in a particular way, they are generally joined by some linker word that carries information about the relationship between the clauses; furthermore, the verbs in all clauses are conjugated normally. In Japanese, all non-main clauses are simply joined, often without any indication of what the relationship is (for matters of time, this is not that unusual; many languages lack such words); as well, the verbs in all but the main clause are deficient - they lack various things like tense, mood, politeness auxiliaries, etc. This is a matter of economy; all that stuff they stick on the verb at the end of the sentence can be rather lengthy. Thus it uses a generic verb form which in some ways resembles the -ing form of English verbs; this form indicates a conjunctive relationship between sentences (note that the literal conjunctive, as in the example a bit above, is actually indicated with a separate form in Japanese, appropriately known as the "conjunctive form/base"; what I've done is merged the two uses).
Here are some examples of things that would use this conjunctive form in Japanese. The first version is how it would be said in Japanese (note that I'm conjugating all verbs here, even though only the last one would be conjugated in Japanese); the second sentences shows how we would typically say the same thing in English.
Simultaneity: "I looked at manga and she looked at novels"/"I looked at manga while/as she looked at novels"
Coincidence: "I went shopping and ran into a friend"/"I ran into a friend when I went shopping"
Sequence: "I got a haircut, [and] went to the bank, and went to the supermarket"/"I got a haircut, went to the bank, [and] then went to the supermarket"
Consequence: "I overslept and was late for class"/"I was late for class because I overslept"
So that's all the conjunctive form is. On an interesting random note, you might notice that in none of those examples does the first version sound unnatural, and might very well be used by native English speakers in addition to the more precise second versions (though of course it would have sounded extremely strange if I had only conjugated the last verb, like Japanese does). This indicates that even in English this kind of vagueness is used; and for that matter, there are ways of indicating some of those relationships explicitly in Japanese, as well - they just aren't always used.
Other times, when creating, the assumptions themselves are fun to play with and observe the results. In Caia, I initially envisioned all of the major words - nouns, adjectives, and verbs - to be nouns. Nouns became adjectives when used with an attributive particle analogous to "having" (e.g. "woman having beauty" vs. "beautiful woman"). Taking an idea from Japanese, nouns became verbs by an auxiliary verb meaning more or less "do" or "make" (e.g. "make cut" vs. "cut"). Unfortunately, I ultimately concluded that verbs had to be separate, due to both practical concerns (specifically, concerns about making thoughts too long) and theoretical concerns (different nouns require different semantics, which would produce inconsistent theoretical behavior).
With another one of my languages, I took a different route, with some very interesting results. As this language is synthetic (unlike Caia, which is strongly analytic and isolating), I had quite a bit more flexibility. This language was actually modeled on the Altaic languages - Japanese, Korean, Mongolian, Turkish, etc.; as such, I suppose I can't claim that I invented this (what I'm getting to), but merely took what existed in Altaic languages and perfected it to a degree that doesn't exist in nature - at least to my knowledge.
The result is a language is which nouns, adjectives, and verbs are, in fact, all verbs; they are conjugated exactly the same, and play the same grammatical role. Even though this particular language is fictional, and I'm not expecting anybody to actually speak it, this idea might show up in other languages of mine (possibly a real one), as I find it exceptionally elegant. However, I do like to refer to "nouns" as substantives, "adjectives" as attributes, and "verbs" as actions; this is done because there are slightly differences in meaning between the noun bases in the three cases.
I explained previously that Japanese verbs had several base forms - usually distinguished by one vowel - which were then agglutinated with other things to form complex verb forms. I won't describe them again, as I use different names for the ones in my language, which might result in confusion. In my language, there are at least five different base forms of each verb: neutral, conclusive, attributive, participial, instantiative, and conjunctive (note that these names are not final, and I'm open to suggestions).
The conclusive form is the same as the Japanese conclusive: it's the main verb of the last clause in a sentence; it indicates the end of the sentence, and often has a number of suffixes indicating various details about the sentence. The attributive form is the same as one of the two uses of the Japanese attributive: it's the verb of a relative clause; adjectives are attached to nouns by forming relative clauses (e.g. "cat that is fat" vs. "fat cat"). The participial form resembles the second use of the Japanese attributive form: it is a noun referring to the act named by the verb (essentially an English gerund or participle, as in "watching FLCL makes me want to kill people"). The instantiative is unique to my language, and refers to an instance of the verb's action; for example, "a run" would be an instance of the verb "run". The conjunctive is similar to the Japanese -te form (and also includes the Japanese conjunctive base); specifically, it is used in verbs not in the last clause of a sentence (I'll come back to that). Finally, the neutral form is used in agglutination, and has something of a flexible, context-dependent meaning.
So, how exactly does this framework allow the grand unification? Let's look at an example of the specific meanings of the different bases for each word type (substantive, attribute, and action). Although I should note that not every base is necessary to unify the three; some are simply part of the bigger picture for the language.
For the substantive "human":
Conclusive form: "is [a] human"
Attributive form: "who is [a] human"
Participial: "being human"
Instantiative: "human"
Conjunctive: "is [a] human, and..."
For the attribute "fat":
Conclusive form: "is fat"
Attributive form: "who is fat"/"fat"
Participial form: "being fat"
Instantiative form: "fat thing/person"
Conjunctive form: "is fat, and..."
For the action "travel":
Conclusive form: "travel"
Attributive form: "who travels"
Participial form: "traveling"
Instantiative form: "journey"
Conjunctive form: "travel, and..."
Thus we are able to use identical conjugation for each type of word, treating the first two as stative verbs and the last as an active verb, in an elegant unified system. The real key to this, I think, was the separation of participial and instantiative forms. Note that not all actions have an instantiative form; it only exists where it makes sense: where something is produced or performed.
Now that I've explained how the unification works, there's just one more loose end to tie up: the meaning of the conjunctive form. This form is somewhat foreign to English speakers, because English only works this way in one of the circumstances this language uses it for (specifically, conjunction - e.g. "Murasaki is seven years old and [is] surprisingly well-spoken").
In English, when we have multiple clauses in a sentence that are related in a particular way, they are generally joined by some linker word that carries information about the relationship between the clauses; furthermore, the verbs in all clauses are conjugated normally. In Japanese, all non-main clauses are simply joined, often without any indication of what the relationship is (for matters of time, this is not that unusual; many languages lack such words); as well, the verbs in all but the main clause are deficient - they lack various things like tense, mood, politeness auxiliaries, etc. This is a matter of economy; all that stuff they stick on the verb at the end of the sentence can be rather lengthy. Thus it uses a generic verb form which in some ways resembles the -ing form of English verbs; this form indicates a conjunctive relationship between sentences (note that the literal conjunctive, as in the example a bit above, is actually indicated with a separate form in Japanese, appropriately known as the "conjunctive form/base"; what I've done is merged the two uses).
Here are some examples of things that would use this conjunctive form in Japanese. The first version is how it would be said in Japanese (note that I'm conjugating all verbs here, even though only the last one would be conjugated in Japanese); the second sentences shows how we would typically say the same thing in English.
Simultaneity: "I looked at manga and she looked at novels"/"I looked at manga while/as she looked at novels"
Coincidence: "I went shopping and ran into a friend"/"I ran into a friend when I went shopping"
Sequence: "I got a haircut, [and] went to the bank, and went to the supermarket"/"I got a haircut, went to the bank, [and] then went to the supermarket"
Consequence: "I overslept and was late for class"/"I was late for class because I overslept"
So that's all the conjunctive form is. On an interesting random note, you might notice that in none of those examples does the first version sound unnatural, and might very well be used by native English speakers in addition to the more precise second versions (though of course it would have sounded extremely strange if I had only conjugated the last verb, like Japanese does). This indicates that even in English this kind of vagueness is used; and for that matter, there are ways of indicating some of those relationships explicitly in Japanese, as well - they just aren't always used.
Monday, April 28, 2008
Random Linguistic Fact of the Day
An affix is something in linguistics which attaches to the word it modifies. For example, the English plural suffix 's'/'es' is an affix, as shown in "books on the shelf". The suffix attaches to the word made plural - 'book'.
A clitic is something that attaches to something other than the word it modifies. An example of this in English is the possessive 's. Adding that to the previous example gives "books on the shelf's". Here, the possessive clitic is attached to 'shelf', even though the word it's actually modifying (the possessor) is 'books'.
This is the technical term for what Trique uses frequently (and I didn't know the name of until now). In Trique, personal and possessive pronouns often become clitics when following certain other types of words. Trique also uses clitic doubling, in some cases.
A clitic is something that attaches to something other than the word it modifies. An example of this in English is the possessive 's. Adding that to the previous example gives "books on the shelf's". Here, the possessive clitic is attached to 'shelf', even though the word it's actually modifying (the possessor) is 'books'.
This is the technical term for what Trique uses frequently (and I didn't know the name of until now). In Trique, personal and possessive pronouns often become clitics when following certain other types of words. Trique also uses clitic doubling, in some cases.
Friday, April 25, 2008
More Random Thoughts
Well, I've had two more random thoughts about English evolution (so far) today.
First, I realized that there is already a mechanism in common use for English to lose the past tense entirely. Figure it out, yet? I just used it. English has come to like to move auxiliary verbs ("do"/"did") before the subject in questions, e.g. "Did you figure it out, yet?" However, it's becoming common use to drop auxiliary verbs in English. This is one such case. As the auxiliary verb, not the main verb, carries the tense, this change leads to a loss of explicitly stated tense.
In the case of questions, the main verb is kept in the infinitive, which is identical in form to the present tense. Languages have a way of evolving based on analogy. It's not impossible that this could be applied to verbs in general, losing the past tense entirely (or at least a distinct form for the past tense; it's possible a periphrastic form, like "do"/"did" would then take it's place in all cases).
The third thought of the day came from me pondering the second one. This modern tendency to drop auxiliary verbs and sometimes the subject is not unique to questions. It's also applied to the perfect (e.g. "I've been" -> "Been") and the progressive (e.g. "I'm thinking" -> "Thinking"). If the latter case became the standard, the effect would basically be the same as with dropping "do"/"did". However, if the former occurred, it could result in the past participle replacing other forms, such as the perfect or past tense.
Here's where my random thought came in: the possibility of replacing the past tense is especially interesting, because it may have happened before in English. If you look back at Old English, you'll notice there are two distinct forms - past tense and past participle - for both strong verbs, which retain this distinction (verbs like "write"/"wrote"/"written"), as well as weak verbs (what became our modern regular verbs like "poke"/"poked"). The past tense for weak verbs, in Old English, was -de. Care to guess what the past participle was? It's -ed, the modern past tense suffix. In summary, the Old English part participle has become the modern English past tense form.
First, I realized that there is already a mechanism in common use for English to lose the past tense entirely. Figure it out, yet? I just used it. English has come to like to move auxiliary verbs ("do"/"did") before the subject in questions, e.g. "Did you figure it out, yet?" However, it's becoming common use to drop auxiliary verbs in English. This is one such case. As the auxiliary verb, not the main verb, carries the tense, this change leads to a loss of explicitly stated tense.
In the case of questions, the main verb is kept in the infinitive, which is identical in form to the present tense. Languages have a way of evolving based on analogy. It's not impossible that this could be applied to verbs in general, losing the past tense entirely (or at least a distinct form for the past tense; it's possible a periphrastic form, like "do"/"did" would then take it's place in all cases).
The third thought of the day came from me pondering the second one. This modern tendency to drop auxiliary verbs and sometimes the subject is not unique to questions. It's also applied to the perfect (e.g. "I've been" -> "Been") and the progressive (e.g. "I'm thinking" -> "Thinking"). If the latter case became the standard, the effect would basically be the same as with dropping "do"/"did". However, if the former occurred, it could result in the past participle replacing other forms, such as the perfect or past tense.
Here's where my random thought came in: the possibility of replacing the past tense is especially interesting, because it may have happened before in English. If you look back at Old English, you'll notice there are two distinct forms - past tense and past participle - for both strong verbs, which retain this distinction (verbs like "write"/"wrote"/"written"), as well as weak verbs (what became our modern regular verbs like "poke"/"poked"). The past tense for weak verbs, in Old English, was -de. Care to guess what the past participle was? It's -ed, the modern past tense suffix. In summary, the Old English part participle has become the modern English past tense form.
Random Thought of the Day
So, I was chatting with friends via instant messages, while doing some other things. I wrote the emote
This got me thinking. While obviously omitting the subject for first-person sentences in spoken English would introduce ambiguity (as unlike in emotes, there isn't any indication that the speaker is referring to themself), in some cases, such as where the subject is the possessor of the direct object, this would not create an appreciable amount of ambiguity. Once convention takes over, you could say "He catches up on massive backlog of anime" (which is currently ungrammatical), and it would be understood that said backlog belonged to the subject. I wonder if we'll see this happen in the future.
For the skeptics, I note that this (not anything having to do with emotes, but rather the omission of the possessive pronoun) has already occurred for some nouns (especially body parts) in Spanish and Italian. For example, taken directly from one of my Spanish books, you would not say "El estudiante levantó su mano"(literally "The student raised their hand"), but simply "El estudiante levantó la mano" (literally "The student raised the hand", which is understood as belonging to the student).
Completely unrelated fact: In Indo-European languages that still have a male/female distinction for nouns, the word for "hand" is female.
*catches up on massive backlog of anime*You might have noticed before how emotes tend to avoid the use of pronouns referring to the subject. If I had to take a guess, I'd say the reason for this is the fact that the sentence is actually first person (the speaker is also the subject), but the verb forms used refer to the third person; thus any pronoun referring to the speaker would seem out of place.
This got me thinking. While obviously omitting the subject for first-person sentences in spoken English would introduce ambiguity (as unlike in emotes, there isn't any indication that the speaker is referring to themself), in some cases, such as where the subject is the possessor of the direct object, this would not create an appreciable amount of ambiguity. Once convention takes over, you could say "He catches up on massive backlog of anime" (which is currently ungrammatical), and it would be understood that said backlog belonged to the subject. I wonder if we'll see this happen in the future.
For the skeptics, I note that this (not anything having to do with emotes, but rather the omission of the possessive pronoun) has already occurred for some nouns (especially body parts) in Spanish and Italian. For example, taken directly from one of my Spanish books, you would not say "El estudiante levantó su mano"(literally "The student raised their hand"), but simply "El estudiante levantó la mano" (literally "The student raised the hand", which is understood as belonging to the student).
Completely unrelated fact: In Indo-European languages that still have a male/female distinction for nouns, the word for "hand" is female.
Wednesday, April 16, 2008
Synthesis
One way of summarizing the characteristics of a given language, as well as to compare one language to another, is the index of synthesis: a measure of the degree of synthesis - the process of combining multiple units of meaning into smaller numbers of words - in a language. Thus, in other words, the index of synthesis is a measure of how much information is contained in each word.
The index of synthesis ranges from isolating to synthetic. Isolating languages contain exactly one unit of meaning per word, and adding additional meaning to some word requires adding other words that modify it. On the other end, a purely synthetic language (not known to exist), would have one word for each entire sentence.
English has been changing in favor of isolation for the last couple thousand years, and Modern English is close to the isolating end of the spectrum, with many words containing single units of meaning. Off the top of my head, I'm going to say English words other than nouns, pronouns, and verbs are isolating; for example, adjectives and prepositions contain only a single unit of meaning in each word. Verbs are becoming increasingly isolating, adding auxiliary verbs to indicate additional meaning to the base verb (e.g. 'should not have seen'). Nouns are also becoming less synthetic, as there are now only two flavors of most nouns: singular objective (e.g. 'cat'), and plural objective/possessive ('cats', 'cat's', and 'cats'', which are all pronounced the same when spoken).*
The romance languages are a little more synthetic than English. Spanish, for example, maintains distinctions between singular and plural, male and female (four forms total) in both nouns and adjectives; for example, 'gordo', 'gorda', 'gordos', and 'gordas' all mean the adjective 'fat', but they indicate singular male, singular female, plural male, and plural female, respectively. Spanish also contains more information in its verbs, chiefly the animacy level and number (singular/plural) of the subject.
Japanese is much further toward the synthetic end. In Japanese, some word types commonly undergo agglutination (agglutination and fusion are the two types of synthesis): the combining of multiple "independent" words into single words, when the context calls for it. For example, the pronoun 'watashi' ('me') can be fused with the suffix '-tachi' to form 'watashitachi' ('us'). Verbs are also fused with auxiliary verbs, commonly meaning things like the passive voice, causation, negation, or politeness; for example, the nine-syllable monstrosity 'kokoruminiawasezu' is formed from either 'kokorumi' ('trial') or 'kokorumiru' ('to test'; I'm not sure if this is the noun or verb, as they're both identical when agglutinated), 'niau' ('to suit; to match; to become; to be like'), '[sa]seru' ('cause to'), and 'zu' ('not'), from the Lord's Prayer meaning 'lead us not into temptation' (that's the popular archaic version of that line, but a more accurate Modern English version would be 'do not test us/our faith', which is closer to the Japanese version). As well, multiple types of words may be fused together, as in 'shakugan', formed from 'shaku' (this appears to be an abbreviation of 'shakuretsu', meaning 'burning') and 'gan' ('eye[s]'), from an anime called Shakugan no Shana ("Shana of the Burning Eyes").
A Turkish example from Wikipedia illustrates extreme levels of synthesis: 'Çekoslovakyalılaştıramadıklarımızdanmışsınız' (18 syllables, if I counted correctly), meaning "You are said to be one of those that we couldn't manage to convert to a Czechoslovak". Some people believe that Turkish is related to Japanese (the Altaic hypothesis), but this has not been conclusively proven.
* While English may have [sometimes large] agglutinated words such as 'antidisestablishmentarianism' (an agglutination of 'anti-', 'dis-', 'establish', '-ment', '-ary', '-an', and '-ism'), these are generally considered separate words, and so are not considered to be a synthesis of smaller words; this classification is supported by the fact that such agglutinations cannot be made freely, but must be words established through common use. In contrast, Japanese verbs may be freely agglutinated with other verbs/words that make sense (e.g. auxiliary verbs), without creating entirely new words.
The index of synthesis ranges from isolating to synthetic. Isolating languages contain exactly one unit of meaning per word, and adding additional meaning to some word requires adding other words that modify it. On the other end, a purely synthetic language (not known to exist), would have one word for each entire sentence.
English has been changing in favor of isolation for the last couple thousand years, and Modern English is close to the isolating end of the spectrum, with many words containing single units of meaning. Off the top of my head, I'm going to say English words other than nouns, pronouns, and verbs are isolating; for example, adjectives and prepositions contain only a single unit of meaning in each word. Verbs are becoming increasingly isolating, adding auxiliary verbs to indicate additional meaning to the base verb (e.g. 'should not have seen'). Nouns are also becoming less synthetic, as there are now only two flavors of most nouns: singular objective (e.g. 'cat'), and plural objective/possessive ('cats', 'cat's', and 'cats'', which are all pronounced the same when spoken).*
The romance languages are a little more synthetic than English. Spanish, for example, maintains distinctions between singular and plural, male and female (four forms total) in both nouns and adjectives; for example, 'gordo', 'gorda', 'gordos', and 'gordas' all mean the adjective 'fat', but they indicate singular male, singular female, plural male, and plural female, respectively. Spanish also contains more information in its verbs, chiefly the animacy level and number (singular/plural) of the subject.
Japanese is much further toward the synthetic end. In Japanese, some word types commonly undergo agglutination (agglutination and fusion are the two types of synthesis): the combining of multiple "independent" words into single words, when the context calls for it. For example, the pronoun 'watashi' ('me') can be fused with the suffix '-tachi' to form 'watashitachi' ('us'). Verbs are also fused with auxiliary verbs, commonly meaning things like the passive voice, causation, negation, or politeness; for example, the nine-syllable monstrosity 'kokoruminiawasezu' is formed from either 'kokorumi' ('trial') or 'kokorumiru' ('to test'; I'm not sure if this is the noun or verb, as they're both identical when agglutinated), 'niau' ('to suit; to match; to become; to be like'), '[sa]seru' ('cause to'), and 'zu' ('not'), from the Lord's Prayer meaning 'lead us not into temptation' (that's the popular archaic version of that line, but a more accurate Modern English version would be 'do not test us/our faith', which is closer to the Japanese version). As well, multiple types of words may be fused together, as in 'shakugan', formed from 'shaku' (this appears to be an abbreviation of 'shakuretsu', meaning 'burning') and 'gan' ('eye[s]'), from an anime called Shakugan no Shana ("Shana of the Burning Eyes").
A Turkish example from Wikipedia illustrates extreme levels of synthesis: 'Çekoslovakyalılaştıramadıklarımızdanmışsınız' (18 syllables, if I counted correctly), meaning "You are said to be one of those that we couldn't manage to convert to a Czechoslovak". Some people believe that Turkish is related to Japanese (the Altaic hypothesis), but this has not been conclusively proven.
* While English may have [sometimes large] agglutinated words such as 'antidisestablishmentarianism' (an agglutination of 'anti-', 'dis-', 'establish', '-ment', '-ary', '-an', and '-ism'), these are generally considered separate words, and so are not considered to be a synthesis of smaller words; this classification is supported by the fact that such agglutinations cannot be made freely, but must be words established through common use. In contrast, Japanese verbs may be freely agglutinated with other verbs/words that make sense (e.g. auxiliary verbs), without creating entirely new words.
Sunday, April 13, 2008
Random Linguistic Fact of the Day
Indo-European languages characteristically have distinct forms of verbs such that the verbs agree with some basic properties of the subject. English has almost entirely lost this property for most verbs (third-person singular being the only form that still agrees with the subject), although 'be', the most irregular verb in English, gives you a little bit of an idea of how things used to work:
I am
You are
He/she/it is
We/y'all/they are
In Spanish, the original (more complicated) Indo-European agreement system still exists. As is tradition for Indo-European languages, Spanish verbs agree in number and person (first, second, third) with the subject. At least, that's what most speakers and Spanish teachers will tell you. Here's the full list of forms for the indicative present:
Yo [I] creo [believe]
Nosotros [we] creemos
Tú [you, familiar] crees
Vosotros [y'all, familiar] creéis
Él [he]/ella [she]/usted [you, polite] cree (Spanish does not have a neuter gender, and inanimate objects are either male or female)
Ellos [male they]/ellas [female they]/ustedes [you, polite plural] creen
Now, there's something really weird about that system; did you see it? Usted/ustedes are second person pronouns, but verb agreement indicates that they're third person. How do we explain that?
This was something that mystified me until a couple years ago, when I bought the linguistics book I use as a reference, and learned about animacy/empathy hierarchies, which I've touched on before, and had one of those 'aha!' moments. The answer is that verbs don't actually agree with the person of the subject, but rather by the empathy level of the subject. It's very common in empathy hierarchies to see "first person > second person > third person" (which makes logical sense), so it isn't surprising that Indo-European verb agreement approximates the three persons.
What's out of place - that a second person pronoun is placed in the same level as third person pronouns - can be explained by noting that you would use 'tú' with friends, while you would typically use 'usted' with people you are less familiar with. You would clearly have greater empathy for your friends than random people you meet on the street. Thus it is not surprising that the familiar pronoun is higher on the empathy hierarchy.
Thus the actual hierarchy looks like this. There are five logical divisions, and they're listed from highest empathy to lowest. These five are then grouped into three discrete empathy levels, indicated by the numbers:
1. First person
2. Second person familiar
3. Second person polite
3. Third person
3. 'Fourth person' (this is a generic pronoun where we would use 'they' or the passive voice in English, not referring to anybody specific; e.g. "Se habla Español" - "Spanish is spoken")
I am
You are
He/she/it is
We/y'all/they are
In Spanish, the original (more complicated) Indo-European agreement system still exists. As is tradition for Indo-European languages, Spanish verbs agree in number and person (first, second, third) with the subject. At least, that's what most speakers and Spanish teachers will tell you. Here's the full list of forms for the indicative present:
Yo [I] creo [believe]
Nosotros [we] creemos
Tú [you, familiar] crees
Vosotros [y'all, familiar] creéis
Él [he]/ella [she]/usted [you, polite] cree (Spanish does not have a neuter gender, and inanimate objects are either male or female)
Ellos [male they]/ellas [female they]/ustedes [you, polite plural] creen
Now, there's something really weird about that system; did you see it? Usted/ustedes are second person pronouns, but verb agreement indicates that they're third person. How do we explain that?
This was something that mystified me until a couple years ago, when I bought the linguistics book I use as a reference, and learned about animacy/empathy hierarchies, which I've touched on before, and had one of those 'aha!' moments. The answer is that verbs don't actually agree with the person of the subject, but rather by the empathy level of the subject. It's very common in empathy hierarchies to see "first person > second person > third person" (which makes logical sense), so it isn't surprising that Indo-European verb agreement approximates the three persons.
What's out of place - that a second person pronoun is placed in the same level as third person pronouns - can be explained by noting that you would use 'tú' with friends, while you would typically use 'usted' with people you are less familiar with. You would clearly have greater empathy for your friends than random people you meet on the street. Thus it is not surprising that the familiar pronoun is higher on the empathy hierarchy.
Thus the actual hierarchy looks like this. There are five logical divisions, and they're listed from highest empathy to lowest. These five are then grouped into three discrete empathy levels, indicated by the numbers:
1. First person
2. Second person familiar
3. Second person polite
3. Third person
3. 'Fourth person' (this is a generic pronoun where we would use 'they' or the passive voice in English, not referring to anybody specific; e.g. "Se habla Español" - "Spanish is spoken")
Subscribe to:
Posts (Atom)