Wednesday, September 9, 2015

National Conference about Khmer Natural Language Processing

The Khmer Language Processing Consortium is happy to announce the Second Annual Conference on Khmer Natural language Processing (KNLP 2015), where all its members and others working in this field bring together their work in an effort to collaboratively advance together towards building practical Natural Language Processing for Khmer. The first annual conference took place in October 2014.
The Khmer Natural Language Processing Consortium, created in 2014, groups universities, NGOs, private companies and researchers interested on accelerating – through close coordination and collaboration – the creation of effective natural language processing tools for Khmer language. These tools will be used to improve access to information and communication in this language.
Check on the website now: http://khmernlp.org/

Call for Papers

The Conference will address a range of critically important issues and themes relating to the Khmer Natural Language Processing community. Plenary speakers include some of the leading thinkers in these areas.
The Khmer Language Processing Consortium is inviting proposals for paper presentations that address Khmer Natural Language Processing in one of the following areas:
• Text and Speech processing
• Optical Character recognition for Khmer and similar complex scripts
• Automatic Translation to and from Khmer.
• Interpreting and generating spoken and written language
• Natural language interfaces and dialogue systems
• Pattern recognition, applied NLP systems
• Cognitive aspects of natural language processing
• Computation aspects of natural language processing
• NLP-based knowledge science and service science
• Corpus and Language Resources
• Corpus-based language modeling
• Tools and resources for natural language processing
• Theoretical and Applied Linguistics, NLP Applications
• Semantics, syntax and lexicon
• Evaluation of natural language systems
• Information retrieval and extraction, text mining
• Human processing of language and speech
• Languages for Disability
• Ontology Engineering
• Phonetics, phonology and morphology
• Pragmatics and discourse

Important Dates

DateDescription
21 September 2015Paper Submission Deadline
21 October 2015Acceptance/Rejection Notification
4 November 2015Final Submission Deadline
4 December 2015Annual Conference

Tuesday, May 26, 2015

KhmerOCR Demo App Released on GitHub

First of all, as I have already stated in my GitHub, do not expect this release app, the full OCR system but it's only my demo at the first sight to answer to my research using Support Vector Machine in 2013 and slightly updated on 2014. Thanks for understanding.

Since I do not commit my time to continue on this topic, I would prefer to publish the demo and soon will make up the source code to public as well.

Currently people are working on TesseractOCR and we are waiting for result, of course some result can be found with the OCR Team at khmerocr.org, please try out and support this team if any.

Here if you're still interesting to see, mine, please download from GitHub: KhmerOCR.NET-App

Friday, November 7, 2014

More Update From Khmer Type: Khmer OCR Accuracy

KhmerType pushes more update on khmerocr.org that works more better with Khmer OS Battambang font with size of 26pt. It's good, I do hope the solution could work good enough for all legacy fonts: Limon, ABC etc. that in 90s there were so many documents written in that fonts.



The target documents, we should focus: Laws, Story books and many other education books in library.



Let's follow his post in his blog:



Khmer Type: Khmer OCR ត្រូវ​បាន​ប៉ុន្មាន % ហើយ?:   ថ្ងៃនេះ ខ្ញុំ​យក​សៀវភៅ "ប្រវត្តិ​ប្រជាជាតិ​ខ្មែរ" ដែល​វាយ​លើង​វិញ​សម្រាប់​ធ្វើ​សៀវភៅ​អេឡិចទ្រូនិច មក​ប្តូរ​ទៅ​ជា​ពុម្ព​អក្សរ ...




Tuesday, October 14, 2014

First online TesseractOCR Engine Based for KhmerOCR

KhmerType just announced his online KhmerOCR implemented with TesseractOCR engine.

This year is the year of Tesseract OCR engine since every where, every researchers are focusing on it in Cambodia. The TesseractOCR is an opensource OCR engine maintained by Google. In few years ago, there are some people had tried to train Khmer characters with the engine since 2009 but the result was not good enough to go.

Today, Danh Hong, a team leader of his OCR project and well known as the Khmer OS fonts designer with Thim Rithy, moonOS (Unix Kernel OS) founder has announced his result with an online tool: khmerocr.org that allows people to try the scanned document to convert into Khmer Unicode text.

Web Base Interface: KhmerOCR.org

Currently he has asked for people to test and report error to improve the system.

I have taken some tests with my tested file that I have used in my previous research, the result needs a lot of time/tasks to improve.

All my below testing cases are using font: Khmer OS Content

Case 1: Real Scanned Document (no much noise), font size: 32pt

The document is written in Khmer OS Content with font size 32pt, scanned on HP Scanjet G3110 with high resolution which is clear enough. The result is not yet good enough.

Case 1 : Scanned Document
Updated 15/10: Danh Hong helps to train the expected document and the result is good. (See in comment)

Case 2: Printed Text Using MS Paint (no noise), font size: 11pt

I tried this serious document since the size is around the use case of people using.
Even it has no noise but the result is not well enough yet.

Case 2: Printed Text with Small font size

Case 3: Printed Text Using MS Paint (no noise), font size: 48pt

With this font size, in my method with printed text could product around 98% of accuracy. Here is also producing good result



These tests are only just one part of the font face and it is also a result for developers to improve.
I believe TesseractOCR would do more better when more training data are made but it would not produce 100% accuracy as people expected. We need more involvement to make this at least 95% of accuracy together for people to use. That's why OCR conference is formed.

Thank to Danh Hong and his team for this public initiation.
We are waiting other people's result as well.

Monday, October 13, 2014

Welcome to the Khmer OCR Conference, 28th October

The event is now announced to public, the Khmer OCR conference which is hosting at conference hall of Ministry of Posts and Telecommunication.

The OCR (Optical Character Recognition) software has been around in the world to convert the printed text on the image, pdf or scanned paper into the computed text or characters.

Khmer OCR topic has been in the researching phrase long time ago with some researchers already but most of the case, each researcher is trying to solve different issue in the OCR technologies such in as in segmentation (line or character separation), recognition or classification etc.

Now the conference is about to focus on producing the software that work for public uses, the invite all related researchers to discuss about different solution and TODO list for the next steps.

The event is for invited person only, please contact the host: research@niptict.edu.kh if you would like to participate.

The event only happens after some discussion and meeting with the team so far.


Tuesday, September 30, 2014

Want To Write Scientific Paper or Article, Start From Here

If you want to start writing that kind of research, you might need to learn about Structure, Format, Content, and Style of a Journal-Style Scientific Paper to get to understand around what you will write.

The following table is short snapshot to help you out, for the detail, please read this article.

Sunday, September 28, 2014

Stay Tune, The OCR Conference is Delay Again

As receiving the update today, the venue informed us that the conference which was first delayed to 1st of October, it has been postponed again to end of October due to the invitation and want to have all potential researchers on board.

The invitation is promised to come by next week, wait and see.