Postediting

Post-editing (or postediting) is the process whereby humans amend machine-generated translation to achieve an acceptable final product. A person who post-edits is called a post-editor. The concept of post-editing is linked to that of pre-editing. In the process of translating a text via machine translation, best results may be gained by pre-editing the source text  for example by applying the principles of controlled language  and then post-editing the machine output. It is distinct from editing, which refers to the process of improving human generated text (a process which is often known as revision in the field of translation). Post-edited text may afterwards be revised to ensure the quality of the language choices or proofread to correct simple mistakes.

Post-editing involves the correction of machine translation output to ensure that it meets a level of quality negotiated in advance between the client and the post-editor. Light post-editing aims at making the output simply understandable; full post-editing at making it also stylistically appropriate. With advances in machine translation full post-editing is becoming an alternative to manual translation. Practically all computer assisted translation (CAT) tools now support post-editing of machine translated output.

Post-editing and machine translation

Machine translation left the labs to start being used for its actual purpose in the late seventies at some big institutions such as the European Commission and the Pan-American Health Organization, and then, later, at some corporations such as Caterpillar and General Motors. First studies on post-editing appeared in the eighties, linked to those implementations.[1] To develop appropriate guidelines and training, members of the Association for Machine Translation in the Americas (AMTA) and the European Association for Machine Translation (EAMT) set a Post-editing Special Interest Group in 1999.[2]

After the nineties, advances in computer power and connectivity sped machine translation development and allowed for its deployment through the web browser, including as a free, useful adjunct to the main search engines (Google Translate, Bing Translator, Yahoo! Babel Fish). A wider acceptance of less than perfect machine translation was accompanied also by a wider acceptance of post-editing. With the demand for localisation of goods and services growing at a pace that could not be met by human translation, not even assisted by translation memory and other translation management technologies, industry bodies such as the Translation Automation Users Society (TAUS) expect machine translation and post-editing to play a much bigger role within the next few years.[3]

The use of Machine Translation suggests sometimes pre-editing.

Light and full post-editing

Studies in the eighties distinguished between degrees of post-editing which, in the context of the European Commission Translation Service, were first defined as conventional and rapid[4] or full and rapid.[5] Light and full post-editing seems the wording most used today.

Light post-editing implies minimal intervention by the post-editor, as strictly required to help the end user make some sense of the text; the expectation is that the client will use it for inbound purposes only, often when the text is needed urgently, or has a short time span.

Full post-editing involves a greater level of intervention to achieve a degree of quality to be negotiated between client and post-editor; the expectation is that the outcome will be a text that is not only understandable but presented in some stylistically appropriate way, so it can be used for assimilation and even for dissemination, for inbound and for outbound purposes.

At the top end of full post-editing there is the expectation of a level of quality which is indistinguishable from that of human translation. The assumption, however, has been that it takes less effort for translators to work directly from the source text than to post-edit the machine generated version. With advances in machine translation, this may be changing. For some language pairs and for some tasks, and with engines that have been customised with domain specific good quality data, some clients are already requesting translators to post-edit instead of translating from scratch, in the belief that they will attain similar quality at a lower cost.

The light/full classification, developed in the nineties when machine translation still came on a CD-ROM, may not suit advances in machine translation at the light post-editing end either. For some language pairs and some tasks, particularly if the source has been pre-edited, raw machine output may be good enough for gisting purposes without requiring subsequent human intervention.

Post-editing efficiency

Post-editing is used when raw machine translation is not good enough and human translation not required. Industry advises post-editing to be used when it can at least double the productivity of manual translation, even fourfold it in the case of light post-editing.

However, post-editing efficiency is difficult to predict. Various studies from both academia and industry have claimed that post-editing is generally faster than translating from scratch, regardless of language pairs or translators' experience.[6] There is, however, no agreement about how much time can be saved through post-editing in practice (if any at all): While the industry reports on time savings around 40%,[7] some academic studies suggest that time savings under actual working conditions are more likely to be between 0–20%. Professionals have also reported negative productivity gains where corrections require more time than to translate from scratch.[8][9]

Post-editing and the language industry

After some thirty years, post-editing is still “a nascent profession”.[10] What the right profile of the post-editor is has not yet been fully studied. Post-editing overlaps with translating and editing, but only partially. Most think the ideal post-editor will be a translator keen to be trained on the specific skills required, but there are some who think a bilingual without a background in translation may be easier to train.[11] Not much is known either on who the actual post-editors are, whether they tend to be professional translators, whether they work mostly as in-house employees or self-employed, and on which conditions. Many professional translators dislike post-editing, among other reasons because it tends to be paid at lower rates than conventional translations, with the International Association of Professional Translators and Interpreters (IAPTI) having been particularly vocal about it.[12]

The quality of machine translation output for post-editing is higher, and therefore requires less post-editing effort, when the machine translation is provided by a neural, vertical or customised machine translation engine. Translation efficiency gains can be measured by tracking time linguists need to correct the machine translation in the same translation environment, such as XTM Cloud[13], a Translation management system and Computer-assisted translation tool, where post-editing times and linguistic quality assessment results of the post-edited texts can be compared.

There are not clear figures on how big the post-editing pie is within the translation industry. A recent survey showed 50% of language service providers offered it, but for 85% of them it accounted less than 10% of their throughput.[14] Memsource, a web-based translation tool, claims over 50 percent of translations between English and Spanish, French and other languages have been done in its platform combining translation memory with machine translation.[15] Post-editing is also being done through translation crowdsourcing portals such as Unbabel which, by November 2014 claimed to have post-edited over 11 million words.[16]

Productivity and volume estimates are, in any case, moving targets since advances in machine translation, in a significant part driven by the post-edited text being fed back into its engines, will mean the more post-editing is done, the higher the quality of machine translation and the more widespread post-editing will become.

gollark: BAN HIM!
gollark: To <#477912057560432680>, then...
gollark: Also, <@333530784495304705> wrote the Xenon HTML/CSS thing. Probably - it is their shop.
gollark: Tomorrow:```Breaking News: Dan200 sues everyone```
gollark: Stick a #licenses channel in?

See also

References

  1. "Senez, Dorothy. "Post-editing service for machine translation users at the European Commission". Translating and the Computer 20. Proceedings from Aslib conference, 12 & 13 November 1998; Vasconcellos, M. and M. Léon (1985). "SPANAM and ENGSPA: Machine Translation at the Pan American Health Organization". Computational Linguistics 11, 122-136". Missing or empty |url= (help)
  2. "Allen, Jeffrey. "Post-editing", in Harold Somers (ed.) (2003). Computers and Translation. A translator's guide. Benjamins: Amsterdam/Philadelphia, p. 312". Missing or empty |url= (help)
  3. "TAUS website".
  4. "Loffler-Laurian, Anne marie. (1986). Post-édition rapide et post-édition conventionelle: deux modalities d'una activité spécifique¨. Multilingua 5, 81-88". Missing or empty |url= (help)
  5. "Wagner, Elisabeth. "Rapid post-editing of Systran" Translating and the Computer 5. Proceedings from Aslib conference, 10 & 11 November 1983". Missing or empty |url= (help)
  6. Green, Spence, Jeffrey Heer, and Christopher D. Manning (2013). "The Efficacy of Human Post-Editing for Language Translation" (PDF). ACM Human Factors in Computing Systems.CS1 maint: uses authors parameter (link)
  7. Plitt, Mirko and Francois Masselot (2010). "A Productivity Test of Statistical Machine Translation Post-Editing in A Typical Localisation Context" (PDF). Prague Bulletin of Mathematical Linguistics. 93: 7–16. doi:10.2478/v10108-010-0010-x.
  8. Marcello Federico, Alessandro Cattelan, and Marco Trombetti (2012). "Measuring user productivity in machine translation enhanced computer assisted translation" (PDF). Proceedings of the Tenth Biennial Conference of the Association for Machine Translation in the Americas (AMTA), San Diego, CA, October 28 – November 1.CS1 maint: uses authors parameter (link)
  9. Läubli, Samuel, Mark Fishel, Gary Massey, Maureen Ehrensberger-Dow, and Martin Volk (2013). "Assessing post-editing efficiency in a realistic translation environment" (PDF). Proceedings of the 2nd Workshop on Post-editing Technology and Practice. pp. 83–91.CS1 maint: uses authors parameter (link)
  10. "TAUS website".
  11. Hutchins, John (1995). "Reflections on the History and present state of machine translation" (PDF).CS1 maint: uses authors parameter (link)
  12. "IAPTI website".
  13. "XTM International official website".
  14. "Postediting in Practice. A TAUS Report" (PDF). March 2010. p. 13.
  15. "Memsource website".
  16. "Unbabel Launches A Human-Edited Machine Translation Service To Help Businesses Go Global, Localize Customer Support".

Further reading

This article is issued from Wikipedia. The text is licensed under Creative Commons - Attribution - Sharealike. Additional terms may apply for the media files.