Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points...

165

Transcript of Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points...

Page 1: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Stylometry and the interplay of title and L1 in thedi�erent annotation layers in the Falko corpus

Felix Golcher & Marc Reznicek

Humboldt-Universität zu Berlin

QITL 4 � 31.03.2011

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 1 / 76

Page 2: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 2 / 76

Page 3: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from second language acquisition research

Learner Corpus Research1 study of learner language

I patternsI developmentI controlling variables

2 and describe the variability between learners and learner subgroups

What measures can help us uncover hidden patterns in learner data?

1 Are learner dependent variables detectable in learner texts?

2 How do those variables a�ect the learner language?

3 How strong is the in�uence of those variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 3 / 76

Page 4: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from second language acquisition research

Learner Corpus Research1 study of learner language

I patternsI developmentI controlling variables

2 and describe the variability between learners and learner subgroups

What measures can help us uncover hidden patterns in learner data?

1 Are learner dependent variables detectable in learner texts?

2 How do those variables a�ect the learner language?

3 How strong is the in�uence of those variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 3 / 76

Page 5: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from second language acquisition research

Learner Corpus Research1 study of learner language

I patternsI developmentI controlling variables

2 and describe the variability between learners and learner subgroups

What measures can help us uncover hidden patterns in learner data?

1 Are learner dependent variables detectable in learner texts?

2 How do those variables a�ect the learner language?

3 How strong is the in�uence of those variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 3 / 76

Page 6: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from stylometry

Stylometry. . .

1 classi�es texts according to non-linguistic variables:

I authorship (who wrote a piece?)I genderI other such variables.

2 and (ideally) tries to �nd out the important linguistic features.

Can we apply this technique to learner data?

1 Can we automatically �detect� the learners L1 from her texts?

2 What kind of variables play a (confounding) role?

3 Can we isolate the in�uence of di�erent variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 4 / 76

Page 7: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from stylometry

Stylometry. . .

1 classi�es texts according to non-linguistic variables:

I authorship (who wrote a piece?)I genderI other such variables.

2 and (ideally) tries to �nd out the important linguistic features.

Can we apply this technique to learner data?

1 Can we automatically �detect� the learners L1 from her texts?

2 What kind of variables play a (confounding) role?

3 Can we isolate the in�uence of di�erent variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 4 / 76

Page 8: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from stylometry

Stylometry. . .

1 classi�es texts according to non-linguistic variables:I authorship (who wrote a piece?)

I genderI other such variables.

2 and (ideally) tries to �nd out the important linguistic features.

Can we apply this technique to learner data?

1 Can we automatically �detect� the learners L1 from her texts?

2 What kind of variables play a (confounding) role?

3 Can we isolate the in�uence of di�erent variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 4 / 76

Page 9: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from stylometry

Stylometry. . .

1 classi�es texts according to non-linguistic variables:I authorship (who wrote a piece?)I gender

I other such variables.

2 and (ideally) tries to �nd out the important linguistic features.

Can we apply this technique to learner data?

1 Can we automatically �detect� the learners L1 from her texts?

2 What kind of variables play a (confounding) role?

3 Can we isolate the in�uence of di�erent variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 4 / 76

Page 10: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from stylometry

Stylometry. . .

1 classi�es texts according to non-linguistic variables:I authorship (who wrote a piece?)I genderI other such variables.

2 and (ideally) tries to �nd out the important linguistic features.

Can we apply this technique to learner data?

1 Can we automatically �detect� the learners L1 from her texts?

2 What kind of variables play a (confounding) role?

3 Can we isolate the in�uence of di�erent variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 4 / 76

Page 11: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from stylometry

Stylometry. . .

1 classi�es texts according to non-linguistic variables:I authorship (who wrote a piece?)I genderI other such variables.

2 and (ideally) tries to �nd out the important linguistic features.

Can we apply this technique to learner data?

1 Can we automatically �detect� the learners L1 from her texts?

2 What kind of variables play a (confounding) role?

3 Can we isolate the in�uence of di�erent variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 4 / 76

Page 12: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Coming from stylometry

Stylometry. . .

1 classi�es texts according to non-linguistic variables:I authorship (who wrote a piece?)I genderI other such variables.

2 and (ideally) tries to �nd out the important linguistic features.

Can we apply this technique to learner data?

1 Can we automatically �detect� the learners L1 from her texts?

2 What kind of variables play a (confounding) role?

3 Can we isolate the in�uence of di�erent variables?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 4 / 76

Page 13: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Converging research questions

1 Can we quantify the in�uence of the learner's L1 onhis/her language use?

2 How do L1 e�ects show on di�erent linguistic levels?

I lexisI syntaxI morphology

3 To what extent do L1-e�ects lead to ungrammatical structures in thelearner language?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 5 / 76

Page 14: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Converging research questions

1 Can we quantify the in�uence of the learner's L1 onhis/her language use?

2 How do L1 e�ects show on di�erent linguistic levels?I lexisI syntaxI morphology

3 To what extent do L1-e�ects lead to ungrammatical structures in thelearner language?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 5 / 76

Page 15: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Research Questions: Joining two points of view

Converging research questions

1 Can we quantify the in�uence of the learner's L1 onhis/her language use?

2 How do L1 e�ects show on di�erent linguistic levels?I lexisI syntaxI morphology

3 To what extent do L1-e�ects lead to ungrammatical structures in thelearner language?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 5 / 76

Page 16: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 6 / 76

Page 17: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 7 / 76

Page 18: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced by

I general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 19: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced by

I general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 20: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced by

I general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 21: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced by

I general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 22: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced by

I general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 23: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced byI general learning principles (developmental factors)

I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 24: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced byI general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 25: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced byI general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 26: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Interlanguage

Interlanguage in second language acquisition

Studying language learners as a group, we assume that

1 learners of a second/foreign language have asystematic internal grammar (interlanguage: IL)

2 IL is di�erent from the native language (L1) grammar

3 IL is di�erent from the target language (TL) grammar (Selinker 1972)

4 IL has been claimed to be in�uenced byI general learning principles (developmental factors)I the structure of the target language

I the learner's L1 (transfer)

I mode of acquisition, teaching method, learning strategies,psycho-typological aspects, etc.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 8 / 76

Page 27: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Transfer

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 9 / 76

Page 28: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Transfer

Transfer as cross-linguistic in�uence

Large discussion about what transfer is(Gass et al. 1983; Dechert et al. 1989; Ellis 2009)

I processing mechanism

I learning strategy

I performance/competencephenomenon

I constrains on hypothesis building

I structural borrowing

I etc.

Working de�nition

Language transfer refers to any instance of learner data where astatistically signi�cant correlation (or probability-based relation) is shownto exist between some feature of the interlanguage and any other languagethat has been previously acquired (see Ellis 2009)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 10 / 76

Page 29: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Transfer

Transfer as cross-linguistic in�uence

Large discussion about what transfer is(Gass et al. 1983; Dechert et al. 1989; Ellis 2009)

I processing mechanism

I learning strategy

I performance/competencephenomenon

I constrains on hypothesis building

I structural borrowing

I etc.

Working de�nition

Language transfer refers to any instance of learner data where astatistically signi�cant correlation (or probability-based relation) is shownto exist between some feature of the interlanguage and any other languagethat has been previously acquired (see Ellis 2009)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 10 / 76

Page 30: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Transfer

Transfer on di�erent linguistic levels

Transfer operates on various linguistics levels.

Many studies have looked at each level independently.

I phonology (Broselow 1992)I morphology (e.g. Dusková 1984; Jarvis 2000)I syntax (e.g. Odlin 1990)I semantics (e.g. Kellermann 1979)I lexicon (e.g. Ringbom 1992)I conceptualization (e.g. Stutterheim 1999, Slabakova 2000)I etc.

relative contributions of L1 on linguistic levels

[We need] �a reliable way to measure the relative contributions ofthe native language to the ease or di�culty learners have with eachsubsystem and, by implication, the total contribution of transfer to the processof second language acquisition.� (Odlin 2003, p. 439)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 11 / 76

Page 31: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Transfer

Transfer on di�erent linguistic levels

Transfer operates on various linguistics levels.

Many studies have looked at each level independently.

I phonology (Broselow 1992)I morphology (e.g. Dusková 1984; Jarvis 2000)I syntax (e.g. Odlin 1990)I semantics (e.g. Kellermann 1979)I lexicon (e.g. Ringbom 1992)I conceptualization (e.g. Stutterheim 1999, Slabakova 2000)I etc.

relative contributions of L1 on linguistic levels

[We need] �a reliable way to measure the relative contributions ofthe native language to the ease or di�culty learners have with eachsubsystem and, by implication, the total contribution of transfer to the processof second language acquisition.� (Odlin 2003, p. 439)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 11 / 76

Page 32: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Transfer

Transfer on di�erent linguistic levels

Transfer operates on various linguistics levels.

Many studies have looked at each level independently.I phonology (Broselow 1992)I morphology (e.g. Dusková 1984; Jarvis 2000)I syntax (e.g. Odlin 1990)I semantics (e.g. Kellermann 1979)I lexicon (e.g. Ringbom 1992)I conceptualization (e.g. Stutterheim 1999, Slabakova 2000)I etc.

relative contributions of L1 on linguistic levels

[We need] �a reliable way to measure the relative contributions ofthe native language to the ease or di�culty learners have with eachsubsystem and, by implication, the total contribution of transfer to the processof second language acquisition.� (Odlin 2003, p. 439)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 11 / 76

Page 33: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Background SLA Transfer

Transfer on di�erent linguistic levels

Transfer operates on various linguistics levels.

Many studies have looked at each level independently.I phonology (Broselow 1992)I morphology (e.g. Dusková 1984; Jarvis 2000)I syntax (e.g. Odlin 1990)I semantics (e.g. Kellermann 1979)I lexicon (e.g. Ringbom 1992)I conceptualization (e.g. Stutterheim 1999, Slabakova 2000)I etc.

relative contributions of L1 on linguistic levels

[We need] �a reliable way to measure the relative contributions ofthe native language to the ease or di�culty learners have with eachsubsystem and, by implication, the total contribution of transfer to the processof second language acquisition.� (Odlin 2003, p. 439)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 11 / 76

Page 34: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Learner Corpus Research on transfer

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 12 / 76

Page 35: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Learner Corpus Research on transfer

Transfer as overuse/underuse

Some studies have studied transfer by comparing frequencies ofPOS-tag n-grams (n < 5, see Aarts et al. 1998; Borin et al. 2004)

POS-tag-chains which show a signi�cant overuse/underuse for a special

L1 or subgroup of L1s indicate transfer e�ects

In a similar approach Zeldes et al. (2008) study L1 independent ILstructures

POS-tag-chains which show a signi�cant overuse/underuse for all L1sindicate L2-structural di�culties

Only short token-based n-grams have been looked at!

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 13 / 76

Page 36: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Learner Corpus Research on transfer

Transfer as overuse/underuse

Some studies have studied transfer by comparing frequencies ofPOS-tag n-grams (n < 5, see Aarts et al. 1998; Borin et al. 2004)

POS-tag-chains which show a signi�cant overuse/underuse for a special

L1 or subgroup of L1s indicate transfer e�ects

In a similar approach Zeldes et al. (2008) study L1 independent ILstructures

POS-tag-chains which show a signi�cant overuse/underuse for all L1sindicate L2-structural di�culties

Only short token-based n-grams have been looked at!

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 13 / 76

Page 37: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Learner Corpus Research on transfer

Transfer as overuse/underuse

Some studies have studied transfer by comparing frequencies ofPOS-tag n-grams (n < 5, see Aarts et al. 1998; Borin et al. 2004)

POS-tag-chains which show a signi�cant overuse/underuse for a special

L1 or subgroup of L1s indicate transfer e�ects

In a similar approach Zeldes et al. (2008) study L1 independent ILstructures

POS-tag-chains which show a signi�cant overuse/underuse for all L1sindicate L2-structural di�culties

Only short token-based n-grams have been looked at!

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 13 / 76

Page 38: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Learner Corpus Research on transfer

Exploiting stylometry to uncover transfer

A set of studies use machine learning techniques to classify the L1 ofthe author of IL-texts (ICLE version 1/2)

Transfer e�ects on classi�cation

If learner text shows special features unique to just one L1-group anddistinct from all other L1s, this must be due to transfer (if all other groupvariables are equally distributed).

Koppel et al. (2003); Koppel et al. (2005); Tsur et al. (2007);JojoWong et al. (2009); Golcher (to appear)

I L1: Bulgarian, Czech, French, Russian, SpanishI L2: EnglishI measures based on: errors, function words, rare POS-bi-grams,

letter-bi-grams (sub-token)

L1 speci�c structures in the IL were strong enough to recover the L1 highlyabove the baseline. ⇒ transfer can be detected by L1-classi�cation.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 14 / 76

Page 39: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Learner Corpus Research on transfer

Exploiting stylometry to uncover transfer

A set of studies use machine learning techniques to classify the L1 ofthe author of IL-texts (ICLE version 1/2)

Transfer e�ects on classi�cation

If learner text shows special features unique to just one L1-group anddistinct from all other L1s, this must be due to transfer (if all other groupvariables are equally distributed).

Koppel et al. (2003); Koppel et al. (2005); Tsur et al. (2007);JojoWong et al. (2009); Golcher (to appear)

I L1: Bulgarian, Czech, French, Russian, SpanishI L2: EnglishI measures based on: errors, function words, rare POS-bi-grams,

letter-bi-grams (sub-token)

L1 speci�c structures in the IL were strong enough to recover the L1 highlyabove the baseline. ⇒ transfer can be detected by L1-classi�cation.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 14 / 76

Page 40: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Learner Corpus Research on transfer

Exploiting stylometry to uncover transfer

A set of studies use machine learning techniques to classify the L1 ofthe author of IL-texts (ICLE version 1/2)

Transfer e�ects on classi�cation

If learner text shows special features unique to just one L1-group anddistinct from all other L1s, this must be due to transfer (if all other groupvariables are equally distributed).

Koppel et al. (2003); Koppel et al. (2005); Tsur et al. (2007);JojoWong et al. (2009); Golcher (to appear)

I L1: Bulgarian, Czech, French, Russian, SpanishI L2: EnglishI measures based on: errors, function words, rare POS-bi-grams,

letter-bi-grams (sub-token)

L1 speci�c structures in the IL were strong enough to recover the L1 highlyabove the baseline. ⇒ transfer can be detected by L1-classi�cation.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 14 / 76

Page 41: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 15 / 76

Page 42: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Road map

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 16 / 76

Page 43: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Road map

From similarity to transfer

We want to classify IL-texts for author's L1:

We de�ne a similarity measure for texts:I A text is a string of characters.I Take two texts A and B, compute a number S from them.I Interprete this number as an indicator for similarity.

Assign a text to the �most similar� L1 (details later!)

a posteriori justi�cation

If the assignments are correct,

⇒ then S is a re�ection of L1 speci�c structures in IL (⇐ transfer).

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 17 / 76

Page 44: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Road map

From similarity to transfer

Transfer on di�erent linguistic levels

L1 classi�cation results based on di�erent linguistic levelsre�ect transfer on that speci�c level

I lemma ⇒ (mainly) transfer on lexical choiceI Part-of-Speech ⇒ (mainly) syntactic transfer

Transfer and grammatical errors

If there is a di�erence between those results for

(a) the learner text(b) a grammatically corrected version of it (target hypothesis)

then this re�ects transfer leading to ungrammatical IL-structures.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 18 / 76

Page 45: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Road map

From similarity to transfer

Transfer on di�erent linguistic levels

L1 classi�cation results based on di�erent linguistic levelsre�ect transfer on that speci�c level

I lemma ⇒ (mainly) transfer on lexical choiceI Part-of-Speech ⇒ (mainly) syntactic transfer

Transfer and grammatical errors

If there is a di�erence between those results for

(a) the learner text(b) a grammatically corrected version of it (target hypothesis)

then this re�ects transfer leading to ungrammatical IL-structures.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 18 / 76

Page 46: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 19 / 76

Page 47: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko Corpus

Falko1 - Error annotated corpus of advanced learners of German (Lüdelinget al. 2008)

written essays & summaries

L2 essay: 122.789 tokens & L1 essay: 68.485 tokens

44 di�erent L1s

I largest: Danish, English, Russian, French, Uzbek, Turkish

4 di�erent essay topics (titles)

I feminism, wages, criminality, university degree

1http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung-en/falko/standardseite-en(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 20 / 76

Page 48: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko Corpus

Falko1 - Error annotated corpus of advanced learners of German (Lüdelinget al. 2008)

written essays & summaries

L2 essay: 122.789 tokens & L1 essay: 68.485 tokens

44 di�erent L1s

I largest: Danish, English, Russian, French, Uzbek, Turkish

4 di�erent essay topics (titles)

I feminism, wages, criminality, university degree

1http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung-en/falko/standardseite-en(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 20 / 76

Page 49: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko Corpus

Falko1 - Error annotated corpus of advanced learners of German (Lüdelinget al. 2008)

written essays & summaries

L2 essay: 122.789 tokens & L1 essay: 68.485 tokens

44 di�erent L1s

I largest: Danish, English, Russian, French, Uzbek, Turkish

4 di�erent essay topics (titles)

I feminism, wages, criminality, university degree

1http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung-en/falko/standardseite-en(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 20 / 76

Page 50: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko Corpus

Falko1 - Error annotated corpus of advanced learners of German (Lüdelinget al. 2008)

written essays & summaries

L2 essay: 122.789 tokens & L1 essay: 68.485 tokens

44 di�erent L1s

I largest: Danish, English, Russian, French, Uzbek, Turkish

4 di�erent essay topics (titles)

I feminism, wages, criminality, university degree

1http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung-en/falko/standardseite-en(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 20 / 76

Page 51: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko Corpus

Falko1 - Error annotated corpus of advanced learners of German (Lüdelinget al. 2008)

written essays & summaries

L2 essay: 122.789 tokens & L1 essay: 68.485 tokens

44 di�erent L1sI largest: Danish, English, Russian, French, Uzbek, Turkish

4 di�erent essay topics (titles)

I feminism, wages, criminality, university degree

1http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung-en/falko/standardseite-en(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 20 / 76

Page 52: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko Corpus

Falko1 - Error annotated corpus of advanced learners of German (Lüdelinget al. 2008)

written essays & summaries

L2 essay: 122.789 tokens & L1 essay: 68.485 tokens

44 di�erent L1sI largest: Danish, English, Russian, French, Uzbek, Turkish

4 di�erent essay topics (titles)

I feminism, wages, criminality, university degree

1http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung-en/falko/standardseite-en(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 20 / 76

Page 53: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko Corpus

Falko1 - Error annotated corpus of advanced learners of German (Lüdelinget al. 2008)

written essays & summaries

L2 essay: 122.789 tokens & L1 essay: 68.485 tokens

44 di�erent L1sI largest: Danish, English, Russian, French, Uzbek, Turkish

4 di�erent essay topics (titles)I feminism, wages, criminality, university degree

1http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung-en/falko/standardseite-en(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 20 / 76

Page 54: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko data subset for classi�cation

Texts included

languages with at least 10 texts

learners with only one L1

Very small data sample

We use only ≈ 66.000 tokens.This is 34% of Falko.

L1 # of texts

German (deu)a 10English (eng) 42Danish (dan) 37French (fra) 14Russian (rus) 10Turkish (tur) 10total 126 texts

acontrol group, excluded if sensible

title texts

�crime� 11 Kriminalität zahlt sich nicht aus.

�feminism� 23 Der Feminismus hat den Interessen der Frauen mehr geschadet als genützt.

�wages� 60 Die �nanzielle Entlohnung eines Menschen sollte dem Beitrag entsprechen,den er/ sie für die Gesellschaft geleistet hat.

�studies� 32 Die meisten Universitätsabschlüsse sind nicht praxisorientiert und bereitendie Studenten nicht auf die wirkliche Welt vor.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 21 / 76

Page 55: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko - 6 representationsWe have 6 representations of each text.

Each representation is de�ned by two variables:

1 Level of linguistic representation:

token original texts:Man denke an den unterschiedlichen Gruppen, die sich für

den Umweltsschutz einsetzen.

POS Part-of-Speech tag sequence (Treetagger1):PIS VVFIN APPR ART ADJA NN $, PRELS PRF APPR

ART NN VVINF $.

lemma lemma sequence:man denken an d unterschiedlich Gruppe , d er|es|sie für d

Umweltsschutz einsetzen .

2 Level of error contamination:

learner The raw learner texts:Man denke an��den unterschiedlichen Gruppen, die [. . . ]

Target hypothesis the grammaticalized version(Reznicek et al. 2010):Man denke an die unterschiedlichen Gruppen, die [. . . ]

1Schmid 1994.(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 22 / 76

Page 56: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko - 6 representationsWe have 6 representations of each text.

Each representation is de�ned by two variables:1 Level of linguistic representation:

token original texts:Man denke an den unterschiedlichen Gruppen, die sich für

den Umweltsschutz einsetzen.

POS Part-of-Speech tag sequence (Treetagger1):PIS VVFIN APPR ART ADJA NN $, PRELS PRF APPR

ART NN VVINF $.

lemma lemma sequence:man denken an d unterschiedlich Gruppe , d er|es|sie für d

Umweltsschutz einsetzen .

2 Level of error contamination:

learner The raw learner texts:Man denke an��den unterschiedlichen Gruppen, die [. . . ]

Target hypothesis the grammaticalized version(Reznicek et al. 2010):Man denke an die unterschiedlichen Gruppen, die [. . . ]

1Schmid 1994.(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 22 / 76

Page 57: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Current study Our data - the Falko corpus

Falko - 6 representationsWe have 6 representations of each text.

Each representation is de�ned by two variables:1 Level of linguistic representation:

token original texts:Man denke an den unterschiedlichen Gruppen, die sich für

den Umweltsschutz einsetzen.

POS Part-of-Speech tag sequence (Treetagger1):PIS VVFIN APPR ART ADJA NN $, PRELS PRF APPR

ART NN VVINF $.

lemma lemma sequence:man denken an d unterschiedlich Gruppe , d er|es|sie für d

Umweltsschutz einsetzen .

2 Level of error contamination:

learner The raw learner texts:Man denke an��den unterschiedlichen Gruppen, die [. . . ]

Target hypothesis the grammaticalized version(Reznicek et al. 2010):Man denke an die unterschiedlichen Gruppen, die [. . . ]

1Schmid 1994.(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 22 / 76

Page 58: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 23 / 76

Page 59: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 60: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings

A = xabay

B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 61: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings

A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 62: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 63: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1

log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 64: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1

log(2 · 1+ 1) = 1.09

ab 1 1

log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 65: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1

log(2 · 1+ 1) = 1.09

ab 1 1

log(1 · 1+ 1) = 0.69

b 1 3

log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 66: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1

log(2 · 1+ 1) = 1.09

ab 1 1

log(1 · 1+ 1) = 0.69

b 1 3

log(1 · 3+ 1) = 1.39

x 1 0

log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 67: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1

log(

2 · 1

+ 1) = 1.09

ab 1 1

log(

1 · 1

+ 1) = 0.69

b 1 3

log(

1 · 3

+ 1) = 1.39

x 1 0

log(

1 · 0

+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 68: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1

+ 1

)

= 1.09

ab 1 1 log(1 · 1

+ 1

)

= 0.69

b 1 3 log(1 · 3

+ 1

)

= 1.39

x 1 0 log(1 · 0

+ 1

)

= 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 69: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1)

= 1.09

ab 1 1 log(1 · 1+ 1)

= 0.69

b 1 3 log(1 · 3+ 1)

= 1.39

x 1 0 log(1 · 0+ 1)

= 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 70: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 71: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =

∑= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 72: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =∑

= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 73: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S explained by example

Two very short texts:

substrings A = xabay B = bcbabd

a 2 1 log(2 · 1+ 1) = 1.09ab 1 1 log(1 · 1+ 1) = 0.69b 1 3 log(1 · 3+ 1) = 1.39x 1 0 log(1 · 0+ 1) = 0

S =∑

= 3.17

an important feature

All substrings of all lengths contribute:

⇒ No maximal length is set (as is the usual praxis).

No other information than (character) string repetitions are used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 24 / 76

Page 74: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

The similarity measure S � basic concept

S as an established stylometric measure

Various stylometric tasks have been investigated with S :

Translationese: Have translations their own �style�?I Studies in Baroni et al. (2006) have been replicated.

Authorship Attribution: Who wrote the federalist papers? (Golcher2007)

I Main stream attribution of disputed essays con�rmed.

Recoverage of L1 in English:I Replication of the mentioned studies Tsur et al. (2007); Koppel et al.

(2005) (Golcher 2007; Golcher to appear)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 25 / 76

Page 75: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 26 / 76

Page 76: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1

A short reminder of the data.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 27 / 76

Page 77: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1

Falko data subset for classi�cation

Texts included

languages with at least 10 texts

learners with only one L1

Very small data sample

We use only ≈ 66.000 tokens.This is 34% of Falko.

L1 # of texts

German (deu)a 10English (eng) 42Danish (dan) 37French (fra) 14Russian (rus) 10Turkish (tur) 10total 126 texts

acontrol group, excluded if sensible

title texts

�crime� 11 Kriminalität zahlt sich nicht aus.

�feminism� 23 Der Feminismus hat den Interessen der Frauen mehr geschadet als genützt.

�wages� 60 Die �nanzielle Entlohnung eines Menschen sollte dem Beitrag entsprechen,den er/ sie für die Gesellschaft geleistet hat.

�studies� 32 Die meisten Universitätsabschlüsse sind nicht praxisorientiert und bereitendie Studenten nicht auf die wirkliche Welt vor.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 28 / 76

Page 78: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1

Remark

There is no signi�cant correlation between the essay title andthe author's L1.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 29 / 76

Page 79: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1

Some details of the classi�cation method

Take one text Ti after another as test text (126 texts).

following steps:1 Compute S(Ti ,Tj) for the remaining 125 training texts (i 6= j)2 Group those S values according to the L1 of those training texts.3 Compute the mean S value SL1 for each L1 group.4 Assign the test text Ti to the L1 group with the highest SL1.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 30 / 76

Page 80: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 31 / 76

Page 81: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

Proof of concept

Expectation 1

The baseline of random assignments is around 126/6 = 21.We expect to be substantially better than this baseline.

Outcome

With the raw learner texts we get 65 correct assignments out of 126 .⇒ There is something meaningful going on.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 32 / 76

Page 82: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

Proof of concept

Expectation 1

The baseline of random assignments is around 126/6 = 21.We expect to be substantially better than this baseline.

Outcome

With the raw learner texts we get 65 correct assignments out of 126 .⇒ There is something meaningful going on.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 32 / 76

Page 83: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

A short reminder of the di�erent representations of Falko.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 33 / 76

Page 84: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

Falko - 6 representationsWe have 6 representations of each text.

Each representation is de�ned by two variables:1 Level of linguistic representation:

token original texts:Man denke an den unterschiedlichen Gruppen, die sich für

den Umweltsschutz einsetzen.

POS Part-of-Speech tag sequence (Treetagger1):PIS VVFIN APPR ART ADJA NN $, PRELS PRF APPR

ART NN VVINF $.

lemma lemma sequence:man denken an d unterschiedlich Gruppe , d er|es|sie für d

Umweltsschutz einsetzen .

2 Level of error contamination:

learner The raw learner texts:Man denke an��den unterschiedlichen Gruppen, die [. . . ]

Target hypothesis the grammaticalized version(Reznicek et al. 2010):Man denke an die unterschiedlichen Gruppen, die [. . . ]

1Schmid 1994.(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 34 / 76

Page 85: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

L1 classi�cation

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 35 / 76

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

59

48

60

54

44

60

Figure: German L1 texts aredisregarded here.

Expectation 2

token representation shows astronger L1 e�ect than lemma.

Because: lemma ignoresmorphology completely.

Outcome: Big surprise

We could not detect amorphology e�ect.

Expectation 3

The grammaticalized target

hypothesis should scoresomewhat lower than thelearner version.

Outcome

True.Grammatical error correctionlowers accuracy consistentlybut only minimally.

That's not bad, but didn't wemiss some source of similarity?

Page 86: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

L1 classi�cation

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 35 / 76

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

59

48

60

54

44

60

Figure: German L1 texts aredisregarded here.

Expectation 2

token representation shows astronger L1 e�ect than lemma.

Because: lemma ignoresmorphology completely.

Outcome: Big surprise

We could not detect amorphology e�ect.

Expectation 3

The grammaticalized target

hypothesis should scoresomewhat lower than thelearner version.

Outcome

True.Grammatical error correctionlowers accuracy consistentlybut only minimally.

That's not bad, but didn't wemiss some source of similarity?

Page 87: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

L1 classi�cation

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 35 / 76

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

59

48

60

54

44

60

Figure: German L1 texts aredisregarded here.

Expectation 2

token representation shows astronger L1 e�ect than lemma.

Because: lemma ignoresmorphology completely.

Outcome: Big surprise

We could not detect amorphology e�ect.

Expectation 3

The grammaticalized target

hypothesis should scoresomewhat lower than thelearner version.

Outcome

True.Grammatical error correctionlowers accuracy consistentlybut only minimally.

That's not bad, but didn't wemiss some source of similarity?

Page 88: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

L1 classi�cation

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 35 / 76

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

59

48

60

54

44

60

Figure: German L1 texts aredisregarded here.

Expectation 2

token representation shows astronger L1 e�ect than lemma.

Because: lemma ignoresmorphology completely.

Outcome: Big surprise

We could not detect amorphology e�ect.

Expectation 3

The grammaticalized target

hypothesis should scoresomewhat lower than thelearner version.

Outcome

True.Grammatical error correctionlowers accuracy consistentlybut only minimally.

That's not bad, but didn't wemiss some source of similarity?

Page 89: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

L1 classi�cation

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 35 / 76

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

59

48

60

54

44

60

Figure: German L1 texts aredisregarded here.

Expectation 2

token representation shows astronger L1 e�ect than lemma.

Because: lemma ignoresmorphology completely.

Outcome: Big surprise

We could not detect amorphology e�ect.

Expectation 3

The grammaticalized target

hypothesis should scoresomewhat lower than thelearner version.

Outcome

True.Grammatical error correctionlowers accuracy consistentlybut only minimally.

That's not bad, but didn't wemiss some source of similarity?

Page 90: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Preliminary results

L1 classi�cation

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 35 / 76

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

59

48

60

54

44

60

Figure: German L1 texts aredisregarded here.

Expectation 2

token representation shows astronger L1 e�ect than lemma.

Because: lemma ignoresmorphology completely.

Outcome: Big surprise

We could not detect amorphology e�ect.

Expectation 3

The grammaticalized target

hypothesis should scoresomewhat lower than thelearner version.

Outcome

True.Grammatical error correctionlowers accuracy consistentlybut only minimally.

That's not bad, but didn't wemiss some source of similarity?

Page 91: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 36 / 76

Page 92: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Another possible in�uence: Content

Until now we ignored the essay title people wrote about.

Obviously, texts about �crime� will share words.

This of course leads to higher S values.

If this title e�ect is larger than the L1 e�ect, the latter will be masked.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 37 / 76

Page 93: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

This is not a newly discovered problem

This issue is well known from stylometric studies(e.g. Baroni et al. 2006).

There, usually �content words� are removed. (or similar).

Rarely any quantitative investigation is tried(but see Clement et al. 2003)

So, how large is this e�ect? That's what we turn to now...

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 38 / 76

Page 94: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

The same thing with another variable

Now we classify the texts according to essay title, instead of L1.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 39 / 76

Page 95: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

classi�cation according to title

Expectation 4

We expect a strong title e�ect in thetoken and lemma representations.

Outcome

nearly perfect.

Expectation 5

We do not expect a title e�ect in thePOS representation.

Outcome: Surprise.

Also from the POS representationwe can recover the essay title.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 40 / 76

Page 96: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

classi�cation according to title

Expectation 4

We expect a strong title e�ect in thetoken and lemma representations.

Outcome

nearly perfect.

Expectation 5

We do not expect a title e�ect in thePOS representation.

Outcome: Surprise.

Also from the POS representationwe can recover the essay title.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 40 / 76

Page 97: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

classi�cation according to title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

111

56

111 111

55

112

Expectation 4

We expect a strong title e�ect in thetoken and lemma representations.

Outcome

nearly perfect.

Expectation 5

We do not expect a title e�ect in thePOS representation.

Outcome: Surprise.

Also from the POS representationwe can recover the essay title.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 40 / 76

Page 98: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

classi�cation according to title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

111

56

111 111

55

112

Expectation 4

We expect a strong title e�ect in thetoken and lemma representations.

Outcome

nearly perfect.

Expectation 5

We do not expect a title e�ect in thePOS representation.

Outcome: Surprise.

Also from the POS representationwe can recover the essay title.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 40 / 76

Page 99: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

classi�cation according to title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

111

56

111 111

55

112

Expectation 4

We expect a strong title e�ect in thetoken and lemma representations.

Outcome

nearly perfect.

Expectation 5

We do not expect a title e�ect in thePOS representation.

Outcome: Surprise.

Also from the POS representationwe can recover the essay title.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 40 / 76

Page 100: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

A simple heuristic for �ltering out essay title

We divide all S(A,B) in two groups:1 A and B have the same title.2 They have not.

We compute the mean of each group.

Each S value is divided by the mean of its group.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 41 / 76

Page 101: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Classi�cation results before averaging out title

Figure: German L1 texts are disregardedhere.

Expectation 6

If we �lter out the essay title,L1-classi�cation improves.

Outcome

True.

Discussion:

1 transfer on lexical choice ismuch stronger than onsyntax. (lemma > POS)

2 We still see no e�ect ofmorphology. (lemma = tok)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 42 / 76

Page 102: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Classi�cation results before averaging out title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

59

48

60

54

44

60

Figure: German L1 texts are disregardedhere.

Expectation 6

If we �lter out the essay title,L1-classi�cation improves.

Outcome

True.

Discussion:

1 transfer on lexical choice ismuch stronger than onsyntax. (lemma > POS)

2 We still see no e�ect ofmorphology. (lemma = tok)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 42 / 76

Page 103: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Classi�cation results after averaging out title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

71

49

70 70

43

71

Figure: German L1 texts are disregardedhere.

Expectation 6

If we �lter out the essay title,L1-classi�cation improves.

Outcome

True.

Discussion:

1 transfer on lexical choice ismuch stronger than onsyntax. (lemma > POS)

2 We still see no e�ect ofmorphology. (lemma = tok)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 42 / 76

Page 104: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Classi�cation results after averaging out title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

71

49

70 70

43

71

Figure: German L1 texts are disregardedhere.

Expectation 6

If we �lter out the essay title,L1-classi�cation improves.

Outcome

True.

Discussion:

1 transfer on lexical choice ismuch stronger than onsyntax. (lemma > POS)

2 We still see no e�ect ofmorphology. (lemma = tok)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 42 / 76

Page 105: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Classi�cation results after averaging out title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

71

49

70 70

43

71

Figure: German L1 texts are disregardedhere.

Expectation 6

If we �lter out the essay title,L1-classi�cation improves.

Outcome

True.

Discussion:

1 transfer on lexical choice ismuch stronger than onsyntax. (lemma > POS)

2 We still see no e�ect ofmorphology. (lemma = tok)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 42 / 76

Page 106: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Classi�cation results after averaging out title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

71

49

70 70

43

71

Figure: German L1 texts are disregardedhere.

Expectation 6

If we �lter out the essay title,L1-classi�cation improves.

Outcome

True.

Discussion:

1 transfer on lexical choice ismuch stronger than onsyntax. (lemma > POS)

2 We still see no e�ect ofmorphology. (lemma = tok)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 42 / 76

Page 107: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Taking the essay title into account

Classi�cation results after averaging out title

Text Target Hypothesis

tokPOSLemma

Cla

ssifi

catio

n ac

cura

cy

0.0

0.2

0.4

0.6

0.8

1.0

baseline (Monte carlo simulated)

71

49

70 70

43

71

Figure: German L1 texts are disregardedhere.

Expectation 6

If we �lter out the essay title,L1-classi�cation improves.

Outcome

True.

Discussion:

1 transfer on lexical choice ismuch stronger than onsyntax. (lemma > POS)

2 We still see no e�ect ofmorphology. (lemma = tok)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 42 / 76

Page 108: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 43 / 76

Page 109: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

Copied material

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 44 / 76

Texts about the same subject will normally share lexical material.We have an additional problem:

The full title we call �feminism� reads as

Der Feminismus hat den Interessen der Frauen mehrgeschadet als genützt.Feminism damaged the interests of the women rather than ithelped them.

Especially learners tend to copy phrases like �den Interessen derFrauen�.

These long shared substrings make unproportional contributions to S .

explosion of substrings

The number of substrings of astring grows quadratically withits length.

Page 110: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

Expectation

Expectation 5

If we remove copied material we improve classi�cation performance.

Outcome

This is indeed the case.

yes, we can remove copied material.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 45 / 76

Page 111: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

De�nition of �copied material�

We use a simple heuristic to identify copied material

De�nition (copied material)

A string in text B is copied from text A, if

. . . it occurs only once in the source text A.

. . . this is true even if we strip n characters at both sides.

Example (set n to 1)

A Do we have beer or do we have wine, Josef?

B Someone must have been telling lies about Josef K.

applying the de�nition:

� Josef� is copied.

� have b� is not (�have� occurs twice in text A)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 46 / 76

Page 112: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

De�nition of �copied material�

We use a simple heuristic to identify copied material

De�nition (copied material)

A string in text B is copied from text A, if

. . . it occurs only once in the source text A.

. . . this is true even if we strip n characters at both sides.

Example (set n to 1)

A Do we have beer or do we have wine, Josef?

B Someone must have been telling lies about Josef K.

applying the de�nition:

� Josef� is copied.

� have b� is not (�have� occurs twice in text A)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 46 / 76

Page 113: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

De�nition of �copied material�

We use a simple heuristic to identify copied material

De�nition (copied material)

A string in text B is copied from text A, if

. . . it occurs only once in the source text A.

. . . this is true even if we strip n characters at both sides.

Example (set n to 1)

A Do we have beer or do we have wine, Josef?

B Someone must have been telling lies about Josef K.

applying the de�nition:

� Josef� is copied.

� have b� is not (�have� occurs twice in text A)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 46 / 76

Page 114: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

De�nition of �copied material�

We use a simple heuristic to identify copied material

De�nition (copied material)

A string in text B is copied from text A, if

. . . it occurs only once in the source text A.

. . . this is true even if we strip n characters at both sides.

Example (set n to 1)

A Do we have beer or do we have wine, Josef?

B Someone must have been telling lies about Josef K.

applying the de�nition:

� Josef� is copied.

� have b� is not (�have� occurs twice in text A)

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 46 / 76

Page 115: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

Example

n = 2

Zum Schluss glaube ich, dass der Feminismus den Interessen derFrauen sehr viel nützen könne, aber es gibt zu viele Leute, die dieKonzepte des Feminismus schaden, wenn sie dem Feminisumusfür falschen Gründen oder in den falschen Situationen nützen.

At the end I think, that feminism could help the interests of thewomen very much, but there are too many people, which harmthem concepts of feminism, if they help femininism for wrongsreasons or in wrong situations.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 47 / 76

Page 116: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

Example

n = 5

Zum Schluss glaube ich, dass der Feminismus den Interessen derFrauen sehr viel nützen könne, aber es gibt zu viele Leute, die dieKonzepte des Feminismus schaden, wenn sie dem Feminisumusfür falschen Gründen oder in den falschen Situationen nützen.

At the end I think, that feminism could help the interests of thewomen very much, but there are too many people, which harmthem concepts of feminism, if they help femininism for wrongsreasons or in wrong situations.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 47 / 76

Page 117: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

Example

n = 10

Zum Schluss glaube ich, dass der Feminismus den Interessen derFrauen sehr viel nützen könne, aber es gibt zu viele Leute, die dieKonzepte des Feminismus schaden, wenn sie dem Feminisumusfür falschen Gründen oder in den falschen Situationen nützen.

At the end I think, that feminism could help the interests of thewomen very much, but there are too many people, which harmthem concepts of feminism, if they help femininism for wrongsreasons or in wrong situations.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 47 / 76

Page 118: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Getting rid of copied material

An optimum for the parameter n

●●

0.0 0.1 0.2 0.3 0.4 0.5

6070

8090

100

110

120

1 n

corr

ectly

iden

tifie

d fil

es

Topic removed?

simply normedTopic averaged out

classified by

TopicLanguage

Less restrictively filtered

not f

ilter

ed a

t all

Removing copied material helpsidentifying L1.

An approximate optimum isn = 10.

Title identi�cation is nothampered.

Filtering more and more datadamps title and L1 e�ect.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 48 / 76

Page 119: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 49 / 76

Page 120: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Compared classi�cation results

020

4060

8010

012

0

Title unaccounted for

Title averaged out

Copies removed

65 = 52 %

74 = 59 %

84 = 67 %

baseline (Monte carlo simulated)

Figure: Surface (token) representation

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 50 / 76

Page 121: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Expectation

Expectation 5

If we remove copied material we improve classi�cation performance.

Outcome

This is indeed the case.

yes, we can remove copied material.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 51 / 76

Page 122: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ● ●

● ●

22 24 17 18 25 20

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

19 4 3 2 9

12 1

2 6 11 8 6 9

1 1 2 8 1 1

1 7 2

2 8

Figure: Raw text.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 123: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ●

● ● ● ●

24 22 24 15 19 22

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 3 8 1 3

13

1 4 13 6 7 11

1 1 3 8 1

9 1

1 9

Figure: title averaged out.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 124: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ●

● ● ●

23 22 29 15 16 21

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 5 6 4

13

1 3 21 4 4 9

1 1 11 1

8 2

1 9

Figure: copied material removed.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 125: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ●

● ● ●

23 22 29 15 16 21

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 5 6 4

13

1 3 21 4 4 9

1 1 11 1

8 2

1 9

Figure: copied material removed.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 126: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ●

● ● ●

23 22 29 15 16 21

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 5 6 4

13

1 3 21 4 4 9

1 1 11 1

8 2

1 9

Figure: copied material removed.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 127: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ●

● ● ●

23 22 29 15 16 21

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 5 6 4

13

1 3 21 4 4 9

1 1 11 1

8 2

1 9

Figure: copied material removed.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 128: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ●

● ● ●

23 22 29 15 16 21

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 5 6 4

13

1 3 21 4 4 9

1 1 11 1

8 2

1 9

Figure: copied material removed.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 129: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ●

● ● ●

23 22 29 15 16 21

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 5 6 4

13

1 3 21 4 4 9

1 1 11 1

8 2

1 9

Figure: copied material removed.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.

I Those were the mostungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 130: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Classi�cation according L1 Summarizing classi�cation (stylometric) results

Distribution of right and wrong classi�cations

● ● ●

● ● ● ●

● ● ●

23 22 29 15 16 21

37

13

42

14

10

10

126

dan deu eng fra rus tur

dan

deu

eng

fra

rus

tur

real L1

classified as

22 5 6 4

13

1 3 21 4 4 9

1 1 11 1

8 2

1 9

Figure: copied material removed.

1 German is detected with 100%accuracy.

I IL has been claimed to bemore variable.(see Romaine 2003)

2 Most classi�cation errors occurfor English learners.

I In�uence of common EnglishL2 on German L3?(see Cook 2003)

3 Turkish behaves a bit erratic.I Those were the most

ungrammatical texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 52 / 76

Page 131: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 53 / 76

Page 132: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

Where to go from here?

Successful classi�cation is a reliable indicator for existing transfer.

but e�ect sizes can't be readily quanti�ed.

The title e�ect seems to be �stronger� than L1.

but how much?⇒ comparison of classi�cation accuracies is rather indirect.

Can we surpass the stylometric classi�cational view?

1 Can we directly quantify the in�uence of title and L1?

2 Can we directly compare them? For di�erent levels of representation?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 54 / 76

Page 133: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

Building a (linear mixed) model

For each S(A,B) we construct two variables:

sameTitle 1 if A and B share its title, 0 otherwise.sameL1 1 if authors of A and B share L1, 0 otherwise.

Now we set up a model

S = α · sameTitle + β · sameL1+<textspeci�c contributions>+ ε

where ε is a normally distributed error term.

This (linear mixed) model is �tted.

The parameters α and β can be compared.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 55 / 76

Page 134: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

Building a (linear mixed) model

For each S(A,B) we construct two variables:

sameTitle 1 if A and B share its title, 0 otherwise.sameL1 1 if authors of A and B share L1, 0 otherwise.

Now we set up a model

S = α · sameTitle + β · sameL1+<textspeci�c contributions>+ ε

where ε is a normally distributed error term.

This (linear mixed) model is �tted.

The parameters α and β can be compared.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 55 / 76

Page 135: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

Building a (linear mixed) model

For each S(A,B) we construct two variables:

sameTitle 1 if A and B share its title, 0 otherwise.sameL1 1 if authors of A and B share L1, 0 otherwise.

Now we set up a model

S = α · sameTitle + β · sameL1+<textspeci�c contributions>+ ε

where ε is a normally distributed error term.

This (linear mixed) model is �tted.

The parameters α and β can be compared.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 55 / 76

Page 136: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

Building a (linear mixed) model

For each S(A,B) we construct two variables:

sameTitle 1 if A and B share its title, 0 otherwise.sameL1 1 if authors of A and B share L1, 0 otherwise.

Now we set up a model

S = α · sameTitle + β · sameL1+<textspeci�c contributions>+ ε

where ε is a normally distributed error term.

This (linear mixed) model is �tted.

The parameters α and β can be compared.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 55 / 76

Page 137: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

Building a (linear mixed) model

For each S(A,B) we construct two variables:

sameTitle 1 if A and B share its title, 0 otherwise.sameL1 1 if authors of A and B share L1, 0 otherwise.

Now we set up a model

S = α · sameTitle + β · sameL1+<textspeci�c contributions>+ ε

where ε is a normally distributed error term.

This (linear mixed) model is �tted.

The parameters α and β can be compared.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 55 / 76

Page 138: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

Building a (linear mixed) model

For each S(A,B) we construct two variables:

sameTitle 1 if A and B share its title, 0 otherwise.sameL1 1 if authors of A and B share L1, 0 otherwise.

Now we set up a model

S = α · sameTitle + β · sameL1+<textspeci�c contributions>+ ε

where ε is a normally distributed error term.

This (linear mixed) model is �tted.

The parameters α and β can be compared.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 55 / 76

Page 139: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Beyond stylometry, beyond classi�cation

The results

●●●

0.1

0.2

0.5

Layer

Rel

ativ

e in

fluen

ce o

f L1

and

topi

c

Text Target Hypothesis

tokPOSLemma

Figure: L1 (β) divided by title (α) e�ect.

observations

1 essay title always stronger thanL1: All points below 1.

2 Again, no di�erence betweentoken and lemma

3 the L1 in�uence in POS is muchmore pronounced.

4 Removing errors (slightly)weakens L1 in�uence.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 56 / 76

Page 140: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

1 Research Questions: Joining two points of view

2 Background SLAInterlanguageTransfer

3 Learner Corpus Research on transfer

4 Current studyRoad mapOur data - the Falko corpus

5 The similarity measure S � basic concept

6 Classi�cation according L1Preliminary resultsTaking the essay title into accountGetting rid of copied materialSummarizing classi�cation (stylometric) results

7 Beyond stylometry, beyond classi�cation

8 Conclusion

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 57 / 76

Page 141: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

What comes out for stylometry

Stylometric L1 classi�cation is rather successful:I Remember how small the data are (66,000 tokens).I The method is simple and intuitive.I Only substrings, but all substrings are used.

We can quantify the e�ects of L1 and title or content.

Removal of1 title in�uence2 copied material

greatly boosts L1 classi�cation.

That is: S e�ectively measures L1 induced similarity of learner texts.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 58 / 76

Page 142: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

What comes out for learner corpus research

1 The presented similarity measure can be used to detect transfer e�ects.

2 The transfer e�ect on lexical choice seems considerably stronger thanon syntax.

3 Morphological transfer seems to play no signi�cant role in our data.

4 The amount of transfer leading to ungrammaticality seems to beminor.

Warning!

Learner corpus studies widely ignore the in�uence of theessay subject (title).

But it's even quite strong on abstract levels such as thePart-of-Speech representation.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 59 / 76

Page 143: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

Open questions

1 Which substrings in which representation contribute most to thetransfer related similarity?

It is possible to scan the texts character by character andcheck which contributes what.

2 What is the role of morphology?

What happens with representation of morphological annotation?

3 Can we extend the classi�cation to typological similar languagegroups?

4 Is it possible to use the same method as an indicator of pro�ciency?

5 How good are the classi�cation results, if all levels are used incombination?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 60 / 76

Page 144: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

Open questions

1 Which substrings in which representation contribute most to thetransfer related similarity?

It is possible to scan the texts character by character andcheck which contributes what.

2 What is the role of morphology?

What happens with representation of morphological annotation?

3 Can we extend the classi�cation to typological similar languagegroups?

4 Is it possible to use the same method as an indicator of pro�ciency?

5 How good are the classi�cation results, if all levels are used incombination?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 60 / 76

Page 145: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

Open questions

1 Which substrings in which representation contribute most to thetransfer related similarity?

It is possible to scan the texts character by character andcheck which contributes what.

2 What is the role of morphology?

What happens with representation of morphological annotation?

3 Can we extend the classi�cation to typological similar languagegroups?

4 Is it possible to use the same method as an indicator of pro�ciency?

5 How good are the classi�cation results, if all levels are used incombination?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 60 / 76

Page 146: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

Open questions

1 Which substrings in which representation contribute most to thetransfer related similarity?

It is possible to scan the texts character by character andcheck which contributes what.

2 What is the role of morphology?

What happens with representation of morphological annotation?

3 Can we extend the classi�cation to typological similar languagegroups?

4 Is it possible to use the same method as an indicator of pro�ciency?

5 How good are the classi�cation results, if all levels are used incombination?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 60 / 76

Page 147: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

Open questions

1 Which substrings in which representation contribute most to thetransfer related similarity?

It is possible to scan the texts character by character andcheck which contributes what.

2 What is the role of morphology?

What happens with representation of morphological annotation?

3 Can we extend the classi�cation to typological similar languagegroups?

4 Is it possible to use the same method as an indicator of pro�ciency?

5 How good are the classi�cation results, if all levels are used incombination?

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 60 / 76

Page 148: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Conclusion

Thank you

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 61 / 76

Page 149: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Literatur I

Aarts, Jan et al. (1998). �Tag sequences in learner corpora: A key tointerlanguage grammar and discourse�. In: Learner English on computer.Ed. by Sylviane Granger. Studies in language and linguistics. London[u.a.]: Longman, pp. 132�141. ISBN: 0-582-29883-0.

Baroni, Marco et al. (Sept. 2006). �A New Approach to the Study ofTranslationese: Machine-learning the Di�erence between Original andTranslated Text�. In: Literary and Linguistic Computing 21.3,pp. 259�274. ISSN: 0268-1145. DOI: 10.1093/llc/fqi039.

Borin, Lars et al. (2004). �New wine in old skins? A corpus investigation ofL1 syntactic transfer in learner language�. In: Corpora and languagelearners. Ed. by Guy Aston et al. Vol. 17. Studies in corpus linguistics.Amsterdam, Philadelphia: John Benjamins, pp. 67�87. ISBN:9027222886.

Broselow, Ellen (1992). �Transfer and Universals in Second LanguageEpenthesis�. In: Language Transfer in Language Learning. Ed. bySusan M. Gass et al. John Benjamins Publishing Company, pp. 71�86.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 62 / 76

Page 150: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Literatur II

Clement, Ross et al. (2003). �Ngram and Bayesian classi�cation ofdocuments for topic and authorship�. In: Literary and LinguisticComputing 18.4, pp. 423�447.

Cook, Vivian James (2003). E�ects of the second language on the �rst.Vol. 3. Second language acquisition. Clevedon: Multilingual Matters.ISBN: 1853596337. URL:http://www.gbv.de/dms/bs/toc/357041879.pdf.

Dechert, Hans Wilhelm et al., eds. (1989). Transfer in language production.Norwood, NJ: Ablex Publ. Corp. ISBN: 0893913995.

Dusková, L. (1984). �Similarity-An aid or hindrance in foreign languagelearning?� In: Folia Linguistica: Acta Societatis Linguisticae Europaeae18, pp. 103�115.

Ellis, Rod, ed. (2009). The study of second language acquisition. Oxfordapplied linguistics. Oxford [u.a.]: Oxford Univ. Press. ISBN: 978 0 19442257 4.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 63 / 76

Page 151: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Literatur IIIGass, Susan M. et al., eds. (1983). Language transfer in language learning.Issues in second language research. Rowley, Mass.: Newbury House Publ.ISBN: 0883773058.

Golcher, Felix (2007). �A new text statistical measure and its application tostylometry�. In: Corpus Linguistics 2007. University of Birmingham.

� (to appear). �Repetitions in Text�. PhD thesis. Humboldt-Universität zuBerlin.

Jarvis, Scot H. (2000). �Morphological Type, spatial reference, andlanguage transfer�. In: Studies in Second Language Acquisition 22,pp. 535�556. ISSN: 0272-2631.

JojoWong, Sze-Meng et al. (2009). �Contrastive Analysis and NativeLanguage Identi�cation�. In: Australasian Language TechnologyAssociation Workshop 2009. Ed. by Luiz Augusto Pizzato et al.,pp. 53�61.

Juola, Patrick (2004). Ad-hoc Authorship Attribution Competition. URL:http://www.mathcs.duq.edu/~juola/authorship_contest.html

(visited on 03/24/2011).(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 64 / 76

Page 152: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Literatur IVKellermann, Eric (1979). �Transfer and non-transfer: where we are now�.In: Studies in Second Language Acquisition 2.1, pp. 37�58. ISSN:0272-2631.

Koppel, Moshe et al. (2003). �Exploiting Stylistic Idiosyncrasies forAuthorship Attribution�. In: Proceedings of IJCAI'03 Workshop onComputational Approaches to Style Analysis and Synthesis, pp. 69�72.

Koppel, Moshe et al. (2005). �Automatically Determining an AnonymousAuthor's Native Language�. In: Intelligence and Security Informatics.Lecture Notes in Computer Science. Springer, pp. 209�217. URL:http://www.springerlink.com/content/rem6vng8r20ebk3q/.

Lüdeling, Anke et al. (2008). �Das Lernerkorpus Falko�. In: Deutsch alsFremdsprache 45.2, pp. 67�73.

Odlin, Terence (1990). �Word order, metalinguistic awareness, andconstraints on foreign language learning�. In: Second LanguageAcquisition/Foreign Language Learning. Ed. by Bill VanPatten et al.Channel View Publications Ltd, pp. 95�117. ISBN: 9780585256801.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 65 / 76

Page 153: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Literatur VOdlin, Terence (2003). �Cross-linguistic In�uence�. In: Handbook onSecond Language Acquisition. Ed. by Catherine Doughty et al. Blackwell,pp. 436�486.

R Development Core Team (2010). R: A Language and Environment forStatistical Computing. ISBN 3-900051-07-0. R Foundation for StatisticalComputing. Vienna, Austria. URL: http://www.R-project.org.

Reznicek, Marc et al. (2010). Das Falko-Handbuch: Korpusaufbau undAnnotationen: Version 1.0. Berlin. URL:http://www.linguistik.hu-berlin.de/institut/professuren/

korpuslinguistik/forschung/falko (visited on 10/12/2010).Ringbom, Håkan (1992). �On L1 Transfer in L2 Comprehension andproduction�. In: Language Learning 42, pp. 85�112. ISSN: 1467-9922.

Romaine, Suzanne (2003). �Variation�. In: The handbook of secondlanguage acquisition. Ed. by Catherine Doughty. Vol. 14. Blackwellhandbooks in linguistics. Malden, MA [u.a.]: Blackwell, pp. 409�435.ISBN: 0-631-21754-1.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 66 / 76

Page 154: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Literatur VI

Schmid, Helmut (1994). �Probabilistic Part-of-Speech Tagging UsingDecision Trees�. In: Proceedings of the International Conference on NewMethods in Language Processing, pp. 44�49. URL:http://www.ims.uni-stuttgart.de/ftp/pub/corpora/tree-

tagger1.pdf.Selinker, Larry (1972). �Interlanguage�. In: International Review of AppliedLinguistics in Language Teaching 10.3, pp. 209�231. ISSN:0019042X/2007/045-045.

Slabakova, Roumyana (2000). �L1 transfer revisited: the L2 acquisition oftelicity marking in English by Spanish and Bulgarian native speakers�. In:Linguistics 38.4, pp. 739�770. ISSN: 0024-3949.

Stutterheim, Christiane v. (1999). �How language speci�c are processes inthe conceptualiser?� In: Represenations and Processes in LanguageProduction. Ed. by Ralf Klabunde et al. DUV, pp. 153�179.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 67 / 76

Page 155: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Literatur VII

Tsur, Oren et al. (2007). �Using Classi�er Features for Studying the E�ectof Native Language Choice of Written Second Language Words�. In:Proceedings of the Workshop on Cognitive Aspects of ComputationalLanguage Acquisition. ACL, pp. 9�16.

Zeldes, Amir et al. (2008). �What's hard? Quantitative evidence fordi�cult constructions in German learner data�. In: Proceedings of QITL3. Helsinki.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 68 / 76

Page 156: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Norming S

9 Norming S

10 Density plots

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 69 / 76

Page 157: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Norming S

An obvious problem

The similarity measure S as a formula

S(A,B) =∑

all substrings s

log(FA(s)FB(s) + 1)

FA(s) � Frequency of substring s in Text A

Longer texts ⇒ more and more frequent substrings.

S grows with text length!

Length dependency not easy to parametrize.

and that would not be the full story...

An working heuristic is applied.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 70 / 76

Page 158: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Norming S

A life example

Ti

t j

1 2 3 4 5 6 7 8

12

34

56

78

Figure: Dark: low S-values; Light:high S-values.

Eight Dutch authorsa.

One training �le / one test �le.

Each training �le compared witheach test �le.

=> Training File 8 is the shortest one.

=> Darkest column.

=> lowest S values.

aJuola 2004.

Simple: Dividing Columns by their mean.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 71 / 76

Page 159: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Norming S

Averaging out single text dependencies

Ti

t j

1 2 3 4 5 6 7 8

12

34

56

78

Ti

t j

1 2 3 4 5 6 7 8

12

34

56

78

This normed version of S is what we really used.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 72 / 76

Page 160: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Density plots

9 Norming S

10 Density plots

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 73 / 76

Page 161: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Density plots

Distribution of S(A,B) valuesGreen: A and B share title or L1

Red: Di�erent title or L1.

Same title or not? Same L1 or not?

S in arbitrary units

Den

sity

same titlenot same title

S in arbitrary units

Den

sity

same L1not same L1

title much stronger than L1.

But similarity due to L1 is what we are interested in.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 74 / 76

Page 162: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Density plots

Distribution of S(A,B) values after averaging out titleAgain: Green: A and B share L1; Red: Di�erent L1.

with title: title e�ect removed:

S in arbitrary units

Den

sity

same L1not same L1

S in arbitrary units

Den

sity

same L1not same L1

The di�erence is much clearer now.

Classi�cation jumps from 65 to 74 correct decisions (out of 126).

Suspiciously stretched right tail.

⇒ To this we turn now.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 75 / 76

Page 163: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Density plots

Distribution of S(A,B) values after averaging out titleAgain: Green: A and B share L1; Red: Di�erent L1.

with title: title e�ect removed:

S in arbitrary units

Den

sity

same L1not same L1

S in arbitrary units

Den

sity

same L1not same L1

The di�erence is much clearer now.

Classi�cation jumps from 65 to 74 correct decisions (out of 126).

Suspiciously stretched right tail.

⇒ To this we turn now.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 75 / 76

Page 164: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Density plots

Distribution of S(A,B) values after averaging out titleAgain: Green: A and B share L1; Red: Di�erent L1.

with title: title e�ect removed:

S in arbitrary units

Den

sity

same L1not same L1

S in arbitrary units

Den

sity

same L1not same L1

The di�erence is much clearer now.

Classi�cation jumps from 65 to 74 correct decisions (out of 126).

Suspiciously stretched right tail. ⇒ To this we turn now.

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 75 / 76

Page 165: Stylometry and the interplay of title and L1 in the ... · Research Questions: Joining wot points of view 1 Research Questions: Joining two points of view 2 Background SLA Interlanguage

Density plots

Density plots after removing copied material

S in arbitrary units

Den

sity

same L1not same L1

S in arbitrary units

Den

sity

same L1not same L1

The right tail is greatly reduced.

Classi�cation results again jump from 74 to 84 correct (from 126).

(Humboldt-Universität zu Berlin) Stylometry & transfer in Falko QITL 4 76 / 76