{"id":2708,"date":"2014-12-16T15:28:27","date_gmt":"2014-12-16T12:28:27","guid":{"rendered":"http:\/\/ssrlab.by\/?p=2708"},"modified":"2018-11-21T14:20:51","modified_gmt":"2018-11-21T11:20:51","slug":"about-nlp","status":"publish","type":"post","link":"https:\/\/ssrlab.by\/en\/2708","title":{"rendered":"All About Natural Language Processing"},"content":{"rendered":"<p>Download (PDF, 467KB)<\/p>\n<p style=\"text-align: justify;\"><em><span style=\"font-size: 14px;\">Natural Language Processing is the machine handling of written and spoken human communications. Methods draw on linguistics and statistics, coupled with machine learning, to model language in the service\u00a0of automation.<\/span><\/em><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">What Good Is NLP for Business?<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">There are myriad applications. Every business process (or personal need) that involves speech or text \u2014 with volume, velocity, or complexity sufficient to push you\u00a0to seek automated assistance \u2014 can benefit from Natural Language Processing.\u00a0Let\u2019s review, systematically, what NLP\u00a0can do for you. Here are 22 facets, with examples that illustrate both implementations and R&amp;D initiatives.\u00a0Let\u2019s start with\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.informationweek.com\/software\/business-intelligence\/consumer-and-enterprise-search-not-an-ex\/20600068\" target=\"_blank\">computing\u2019s second-oldest application<\/a><\/span>, search, and then explore NLP uses from everyday to analytical to unusual.<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Information Extraction and Search<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">If all the world\u2019s information were neatly binned in database fields, we wouldn\u2019t need search. Information retrieval would be nothing more than queries. But instead, notionally,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/clarabridge.com\/default.aspx?tabid=137&amp;ModuleID=635&amp;ArticleID=551\" target=\"_blank\">80 percent of business-relevant information originates in unstructured form<\/a><\/span>, primarily text. The vast majority of that text is \u201cnatural language\u201d (as opposed to formal language, found for instance in a computer programming or algebraic equation).\u00a0Google and Bing and other search systems use NLP to <em>extract terms from text (#1)<\/em> to populate their indexes and to <em>parse search queries\u00a0(#2)<\/em>. Those terms may include \u201cnamed entities\u201d such as people, companies, brands, ticker symbols, and places. Other features of interest may include dates, addresses, URLs, and the like; NLP will automate<em>\u00a0extraction of\u00a0pattern-identified information (#3)<\/em> and <em>extraction of attributes associated with terms (#4)<\/em>\u00a0whether factual or subjective:\u00a0expensive\u00a0watch,\u00a0black\u00a0car,\u00a04.6 kg\u00a0fish.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">The more advanced engines apply NLP to <em>identify relationships\u00a0(#5)<\/em>\u00a0(\u201cthis is a that\u201d) in order to build their\u00a0knowledge graphs. NLP\u00a0feeds the computational knowledge engines behind\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.apple.com\/ios\/siri\/\" target=\"_blank\">Apple Siri<\/a><\/span>,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.wolframalpha.com\/\" target=\"_blank\">Wolfram Alpha<\/a><\/span>, and\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.google.com\/landing\/now\/\" target=\"_blank\">Google Now<\/a><\/span>\u00a0as well as resources for your own lexical analyses such as\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.lexalytics.com\/technical-info\/concept-matrix-semantic-analysis\" target=\"_blank\">Lexaltyics\u2019 Concept Matrix<\/a><\/span>, built via NLP application to the Wikipedia dataset to identify \u201cconcept topics\u201d and \u201cfacets\u201d as well as associated sentiment. According to Lexalytics CEO Jeff Catlin, \u201cthese features allow users to easily build classifiers for very broad\u00a0topics as well as roll-up opinions into buckets of similarity.\u201d\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.pingar.com\/introducing-pingar-taxonomy-generator\" target=\"_blank\">Pingar\u2019s Taxonomy Generator<\/a><\/span>\u00a0is another take on the same idea: Use NLP methods to build a knowledge structure for later application to search, classification, and other business information-management needs.<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Concepts, Topics, Sentiment, and Similarity, Plus Notes on Methods<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">\u201cBuckets of similarity\u201d: Those would be categories determined by an analyst or via statistical clustering. Classification is the act of placing cases into categories \u00a0based on attributes or into clusters based on best fit. <em>Classification (#6)<\/em> is part of the NLP task, whether it involves grouping terms or documents. One variety of term grouping involves creating conceptual classes, for instance \u201cvehicle manufacturers\u201d from Fiat, Ford, General Motors, Nissan, Toyota, etc. Another variety involves coreference \u2014 multiple ways of referring to a given thing;\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.whitehouse.gov\/administration\/president-obama\" target=\"_blank\">to illustrate<\/a><\/span>, \u201cBarack H. Obama\u00a0is the\u00a0<span style=\"text-decoration: underline;\">44th President of the United States<\/span>.\u00a0<span style=\"text-decoration: underline;\">His<\/span>\u00a0story is the American story\u2026\u00a0<span style=\"text-decoration: underline;\">President Obama<\/span>\u00a0was born in Hawaii\u201d refers to a given person in four underlined ways, one of them via a pronoun (\u201chis\u201d) that refers to that person only in context.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Want to see real-world entity extraction and coreference? Try Language Computer Corporation\u2019s\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/demo.languagecomputer.com\/cicerolite\/\" target=\"_blank\">Cicero system demo<\/a><\/span>. Process the Web page where I found the above lines,\u00a0http:\/\/www.whitehouse.gov\/administration\/president-obama. Click on one of the \u201che\u201d or \u201chis\u201d occurrences in the marked-up text and you\u2019ll see that these pronouns have been correctly resolved to \u201cPresident Obama.\u201d<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">I suppose I\u2019ll grant numbers to <em>concept extraction (#7)<\/em> and to <em>topic extraction (#8)<\/em>, that is, to information extraction (per the previous subsection) that involves abstraction.\u00a0Sentiment is also abstract, although <em>sentiment analysis (#9)<\/em> can be characterized (in a very simplistic way) as simply another classification problem, whether involving the usual positive\/negative\/neutral categories, more nuanced emotion categories (e.g., angry, happy, sad), or intent signals (e.g., to buy, sell, renew, cancel).\u00a0Visit the Web site of\u00a0text-analytics mavens\u00a0Daedalus for an online\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/showroom.daedalus.es\/en\/language-technologies\/opinioncl\/opinioncl.php\" target=\"_blank\">sentiment classification<\/a><\/span>\u00a0demo. The\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/nerily.com\/#demo\" target=\"_blank\">Nerily online demo<\/a><\/span>\u00a0will extract a variety of other text features.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Sentiment analysis and opinion mining are central topics for me. I\u2019ve written a lot about them, and I organize a twice-yearly conference, the\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/sentimentsymposium.com\/\" target=\"_blank\">Sentiment Analysis Symposium<\/a><\/span>, next up May 8, 2013 in New York, preceded on May 7 by an optional, half-day\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/sentimentsymposium.com\/research.html\" target=\"_blank\">Research &amp; Innovation<\/a><\/span>\u00a0session and an\u00a0optional, half-day\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/sentimentsymposium.com\/tutorial.html\" target=\"_blank\">Practical Sentiment Analysis<\/a><\/span>\u00a0tutorial. A disclosure: Lexalytics, cited above and later in this article, is a sponsor.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">All of this information extraction is what makes NLP a key asset for text analytics, which\u00a0models\u00a0and\u00a0structures\u00a0the\u00a0information content\u00a0of\u00a0textual sources\u00a0for\u00a0business intelligence,\u00a0exploratory data analysis,\u00a0research, or\u00a0investigation. (That\u2019s a definition I wrote back in 2007, in a TechWeb article, that made its way to Wikipedia.)<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">I\u2019ll digress to explain that you can automate human handling of many Natural Language Processing tasks, via a crowdsourcing using\u00a0<span style=\"text-decoration: underline;\"><a href=\"https:\/\/senti.crowdflower.com\/senti\" target=\"_blank\">CrowdFlower<\/a><\/span>\u00a0for \u201chuman-powered sentiment analysis\u201d and other systems built on platforms such as\u00a0<span style=\"text-decoration: underline;\"><a href=\"https:\/\/requester.mturk.com\/tour\/sentiment\" target=\"_blank\">Amazon Mechanical Turk<\/a><\/span>.\u00a0Also, you can also extract sentiment and other information by analyzing non-textual sources that range from transaction records to images and speech.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">We\u2019ll get back to speech bit later. For now, I\u2019ll cite one last function related to classification and similarity, and then\u00a0let\u2019s change tacks. That last for-now function is\u00a0<em>plagiarism detection (#10)<\/em>, essentially passage-similarity evaluation across retrieved text, as explained on the\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.uni-weimar.de\/medien\/webis\/research\/events\/pan-13\/pan13-web\/plagiarism-detection.html\" target=\"_blank\">PAN-13 conference site<\/a><\/span>, with a bit of data and source code to get the Python programmers among you started. (PAN =\u00a0Plagiarism Analysis,\u00a0Authorship Identification, and Near-Duplicate Detection. I guess PAAINDD is kind of awkward as a an acronym.)<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Spelling, Grammar, and Style<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Want to right gud? Lucky for you: NLP is built into your favorite word-processing software. <em>Spell check (#11)<\/em> is NLP at its most basic. Spell check will flag a word that\u2019s not in the dictionary and maybe suggest corrections.\u00a0If you have ever written a document with Microsoft Word (or OpenOffice, Google Docs or any of countless other authoring environments), you\u2019ve seen a spelling checker. But spell check won\u2019t identify the two errors in \u201cI went their at tree o\u2019clock.\u201d Try that sentence in\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.jspell.com\/public-spell-checker.html\" target=\"_blank\">JSpell<\/a><\/span>\u00a0or at\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.spellcheck.net\/\" target=\"_blank\">SpellCheck.net<\/a><\/span>\u00a0as proof. For syntactic errors, you need a grammar checker. And how does a machine check grammar?<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">A linguistic approach to grammar checking might involve resolving parts of speech, via <em>sentence diagramming (#12)<\/em>,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/1aiway.com\/nlp4net\/services\/enparser\/default.aspx\" target=\"_blank\">illustrated here<\/a><\/span>, <em>part-of-speech tagging (#13)<\/em>, as seen in a\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/cogcomp.cs.illinois.edu\/demo\/pos\/?id=4\" target=\"_blank\">Univ. of Illinois demo system<\/a><\/span>, or via <em>study of syntactic relations (#14)<\/em>, \u00e0 la this<span style=\"text-decoration: underline;\">\u00a0<a href=\"http:\/\/www.connexor.com\/nlplib\/?q=demo\/syntax\" target=\"_blank\">Connexor demo<\/a><\/span>. (I\u2019m a bit behind myself, actually. Syntactic parsing is one method of discerning relationships among entities, my <em>#5<\/em>, above.)<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">What do some of the tools out there think of my writing? I pasted the three-sentence paragraph above into one. It found \u201c3 critical writing issues\u201d \u2014 two alleged spelling errors and an accusation of wordiness \u2014 and said my writing is \u201cweak, needs revision.\u201d\u00a0(Free access doesn\u2019t provide detail so I\u2019ll withhold the tool\u2019s name.) Try some others:\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.languagetool.org\/\" target=\"_blank\">LanguageTool<\/a><\/span>\u00a0open source proofreading software (which I didn\u2019t find particularly useful, but you might) and\u00a0<a href=\"http:\/\/www.mystilus.com\/Interactive_check\" target=\"_blank\">Stilus<\/a>\u00a0from my friends at\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.daedalus.es\/en\/\" target=\"_blank\">Daedalus<\/a><\/span>.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Two more varieties of stylistic analysis to mention: Lymbix analyzes e-mail sentiment via the\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/tonecheck.com\/\" target=\"_blank\">ToneCheck<\/a><\/span>\u00a0tool, and automated social-comment moderation is another interesting application, although I\u2019ve been unable to identify an independent provider comparable to\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/techcrunch.com\/2010\/06\/17\/huffington-post-buys-adapative-semantics\/\" target=\"_blank\">Adaptive Semantics<\/a><\/span>, which the Huffington Post bought back in 2010.<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Summarization and Translation<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\"><em>Text summarization (#15)<\/em> is the first of several NLP functions I\u2019ll cite that involve both natural-language understanding (covered in <em>#1<\/em> through<em> #14<\/em>) and natural-language generation. A summarizer has to understand the source text sufficiently to generate a shortened version that is faithful to the content and purpose of the original. Abstracting is a related function<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Visionary researcher Hans Peter Luhn described an approach to automatic text abstracting in has April, 1958 IBM Journal paper,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/altaplana.com\/ibm-luhn58-LiteratureAbstracts.pdf\" target=\"_blank\">The Automatic Creation of Literature Abstracts<\/a><\/span>: \u201cStatistical information derived from word frequency and\u00a0distribution is used by the machine to compute a relative measure of significance, first for individual words\u00a0and then for sentences. Sentences scoring highest in significance are extracted and printed out to become\u00a0the \u2018auto-abstract\u2019.\u201d<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Developer\u00a0Andreas Gohr, at his\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.splitbrain.org\/services\/ots\" target=\"_blank\">SplitBrains.org site<\/a><\/span>, provides a Web interface to\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/cs.haifa.ac.il\/~rotemn\/\">Nadav Rotem<\/a>\u2018s<\/span> open source\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/libots.sourceforge.net\/\" target=\"_blank\">Open Text Summarizer<\/a><\/span>\u00a0code. Try it!<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\"><em>Machine translation (#16)<\/em> is a wonderful NLP application. It doesn\u2019t require explanation; I\u2019ll just point you to\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/translate.google.com\/\" target=\"_blank\">Google translate<\/a><\/span>\u00a0where you can try it yourself. Note the automatic <em>language identification (#17)<\/em> feature.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Translation involves more than just rendering words from one language to another. Each language has its own syntax and idioms. A translator, whether a human or a machine, needs to make sense of the text provided and to make sense in the destination language. That is, like summarization, machine translation involves natural-language generation. So does the next example.<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Question Answering<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\"><a href=\"http:\/\/www-03.ibm.com\/innovation\/us\/watson\/\" target=\"_blank\">IBM Watson<\/a>\u00a0is the most prominent example of <em>question answering (#17)<\/em> at work: Information retrieval that produces usable guidance \u2014 situationally-relevant facts, in a form that reflects question context \u2014 to respond to a query.\u00a0When Watson played Jeopardy, it formulated responses as questions; very different from how it will respond to medical-diagnostic challenges.\u00a0I\u2019ll point you to an academic illustration,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/start.csail.mit.edu\/\" target=\"_blank\">START<\/a><\/span>\u00a0from Boris Katz and associates at MIT, and also refer you to the\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.easyask.com\/products\/natural-language\/\" target=\"_blank\">EasyAsk<\/a><\/span>\u00a0and\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.inbenta.com\/en\/blog\/item\/383-freedom-of-expression-for-the-customers.html\" target=\"_blank\">Inbenta<\/a>\u00a0<\/span>Web sites for explanations how Q-A can work in general business contexts.<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Speech Recognition<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Let\u2019s recognize that speech is natural language too, and cite <em>speech recognition (#18)<\/em> and speech generation or <em>synthesis (#19)<\/em> as two more NLP functions.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Speech is more than just spoken text. It conveys genre, sentiment, mood, and emotion, detectable from word and sentence inflection (an interrogatory sentence \u2014 a question \u2014 is inflected up at the end) and from changes in speech volume and rapidity and other indicators. Check out a recent\u00a0IEEE Spectrum\u00a0podcast,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/spectrum.ieee.org\/podcast\/computing\/software\/teaching-computers-to-hear-emotions\/\" target=\"_blank\">Teaching Computers to Hear Emotions<\/a><\/span>, an interview with University of Rochester\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.ece.rochester.edu\/~wheinzel\/\" target=\"_blank\">Professor\u00a0Wendi Heinzelman<\/a><\/span>, and you\u2019ll hear what I mean.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">You don\u2019t have to render the spoken word as text in order to make analytical use of it, also <em>speech transcription (#20)<\/em> certainly counts as an NLP function. Plenty of academic and industrial work has been done on phonological analysis, which examines sounds and sound patterns, and there are industrial systems that perform <em>voice search (#21)<\/em> on phonemes and patterns. If you\u2019d like to see phonetic transcription in action, check out\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/showroom.daedalus.es\/en\/language-technologies\/phonetictrans\/phonetictrans.php\" target=\"_blank\">Daedalus\u2019s online demo<\/a><\/span>.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">On the flip side, <em>text-to-speech (#22)<\/em> \u2014 having the machine read to you with properly accented pronunciation, inflection, pacing, etc. \u2014 is another bit of NLP. Ivona,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/phx.corporate-ir.net\/phoenix.zhtml?c=176060&amp;p=irol-newsArticle&amp;ID=1777642&amp;highlight=\" target=\"_blank\">recently acquired by Amazon<\/a><\/span>\u00a0\u2014 the software is already used on the Kindle Fire \u2014 has a cool\u00a0<a href=\"http:\/\/www.ivona.com\/us\/\" target=\"_blank\">online demo<\/a>\u00a0that will read for you in a wide variety of languages and accents.<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Building Blocks<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Finally, a note on tools, on the bits and pieces of code you can apply to hobble together your own solution, and on learning more. A disclaimer however: I didn\u2019t intend, in this article, to systematically catalog available software and services, open source or other.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Cognitive linguist Christopher Phipps observes,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/thelousylinguist.blogspot.com\/2013\/01\/free-online-nlp-resources-nltk-still.html\" target=\"_blank\">in his Lousy Linguist blog<\/a><\/span>,\u00a0\u201dluckily, the NLP field has matured into an open access friendly crowd, so there are lots of resources freely available.\u201d Phipps focuses on text understanding, and for that, there\u2019s no better catalog than\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www-nlp.stanford.edu\/links\/statnlp.html\" target=\"_blank\">Stanford University NLP\u2019s page<\/a><\/span>, \u201cStatistical natural language processing and corpus-based computational linguistics: An annotated list of resources,\u201d although as\u00a0Phipps cautions, it\u2019s not for newbies. I list a number of open source tools in a last-year blog article,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/breakthroughanalysis.com\/2012\/01\/08\/what-are-the-most-powerful-open-source-sentiment-analysis-tools\/\" target=\"_blank\">What are the most powerful open-source sentiment-analysis\u00a0tools?<\/a><\/span>\u00a0Two I didn\u2019t include there, because they\u2019re not optimized for sentiment (that article\u2019s topic), are\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/opennlp.apache.org\/\" target=\"_blank\">Apache OpenNLP<\/a><\/span>. and the\u00a0<a href=\"http:\/\/mallet.cs.umass.edu\/index.php\" target=\"_blank\">Mallet<\/a>\u00a0machine-learning toolkit.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">You also have at your disposal a plethora of service offerings that implement NLP, invokable via online APIs, most with free for either trials or limited use. Off the top of my head, there are:\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.alchemyapi.com\/\" target=\"_blank\">AlchemyAPI<\/a>,\u00a0<a href=\"https:\/\/store.apicultur.com\/\" target=\"_blank\">Apicultur<\/a>,\u00a0<a href=\"http:\/\/www.bitext.com\/bitext-api-2.html\" target=\"_blank\">Bitext<\/a>,\u00a0<a href=\"http:\/\/www.clarabridge.com\/Product\/HowClarabridgeWorks\/Implement\/ClarabridgeAPI.aspx\" target=\"_blank\">Clarabridge<\/a>,\u00a0<a href=\"https:\/\/developer.conveyapi.com\/\" target=\"_blank\">ConveyAPI<\/a>,\u00a0<a href=\"http:\/\/www.openamplify.com\/insights\" target=\"_blank\">OpenAmplify<\/a>,\u00a0<a href=\"http:\/\/www.pingar.com\/get-the-api\/\" target=\"_blank\">Pingar<\/a>,\u00a0<a href=\"https:\/\/saplo.com\/\" target=\"_blank\">Saplo<\/a>,\u00a0<a href=\"https:\/\/semantria.com\/\" target=\"_blank\">Semantria<\/a><\/span>\u00a0(backed by the\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.lexalytics.com\/technical-info\/salience-engine-for-text-analysis\" target=\"_blank\">Lexalytics Salience engine<\/a><\/span>), and\u00a0<span style=\"text-decoration: underline;\"><a href=\"https:\/\/www.viralheat.com\/\" target=\"_blank\">Viralheat<\/a><\/span>.\u00a0<span style=\"text-decoration: underline;\"><a href=\"https:\/\/www.mashape.com\/search?query=text\" target=\"_blank\">Mashape lists many more<\/a><\/span>. Capabilities, quality, and cost vary widely. Some do only entity or sentiment tagging while others do more-elemental text analysis. The Apicultur service and Jacob Perkins\u2019 Web API for Python NLTK, at\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/text-processing.com\/\" target=\"_blank\">text-processing.com<\/a><\/span>, are examples of the latter.\u00a0I\u2019ll withhold detail and judgments but maybe write them out in an article at some point, and I\u2019m also not going to write now about install-yourself software options, which aren\u2019t as easy to simply try as a Web API.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">As for learning more, other than via do-it-yourself, what better way than through an online course?\u00a0<span style=\"text-decoration: underline;\"><a href=\"https:\/\/www.coursera.org\/course\/nlangp\" target=\"_blank\">Coursera has one going, taught by Michael Collins of Columbia University, and the videos and lecture materials from\u00a0<\/a><a href=\"http:\/\/see.stanford.edu\/see\/courseInfo.aspx?coll=63480b48-8819-4efd-8412-263f1a472f5a\" target=\"_blank\">Christopher Manning\u2019s popular Stanford University course<\/a><\/span>\u00a0are available online. A third option is the\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.statistics.com\/language\/\" target=\"_blank\">Statistics.com course<\/a><\/span>\u00a0taught by Dr. Nitin Indurkhya, planned for a July 19 start.<\/span><\/p>\n<h3 style=\"text-align: justify;\"><em><strong><span style=\"font-size: 14px;\">Now That You Know<\/span><\/strong><\/em><\/h3>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">In business contexts, you\u2019re most often\u00a0going to apply NLP\u00a0in conjunction with collection, integration, and analysis of disparate forms of online, social, and enterprise data. All that text-and speech-extractable\u00a0goodness I\u2019ve been discussing: In today\u2019s world of heterogeneous big data, it doesn\u2019t stand on its own. This statement is true for business analytics \u2013\u00a0you get lift by applying and integrating\u00a0an appropriate variety of methods and data \u2014 and it\u2019s also true for activities that seemingly don\u2019t involve non-textual or non-speech sources, for activities such as Web search. Even in those latter cases, smart,\u00a0<span style=\"text-decoration: underline;\"><a href=\"http:\/\/www.allanalytics.com\/author.asp?section_id=1415&amp;doc_id=252614\" target=\"_blank\">sense-making engines<\/a><\/span>\u00a0take into account your profile, location, past online\/on-social activities, and social connections, in conjunction with NLP, to provide the best situationally-relevant results.<\/span><\/p>\n<p style=\"text-align: justify;\"><span style=\"font-size: 14px;\">Natural-language processing can do many things for you. It\u2019s an essential tool for leading-edge analytics. Understanding is just the start.<\/span><\/p>\n<p style=\"text-align: justify;\"><strong>Source:<\/strong>\u00a0<a href=\"http:\/\/breakthroughanalysis.com\/2013\/03\/04\/all-about-natural-language-processing\/\">http:\/\/breakthroughanalysis.com\/2013\/03\/04\/all-about-natural-language-processing\/<\/a><\/p>\n<p><span style=\"font-size: 14px;\"><iframe src=\"\/\/docs.google.com\/viewer?url=http%3A%2F%2Fssrlab.by%2Fwp-content%2Fuploads%2F2014%2F12%2Feng_All_About_Natural_Language_Processing.pdf&hl=en_US&embedded=true\" class=\"gde-frame\" style=\"width:100%; height:500px; border: none;\" scrolling=\"no\"><\/iframe>\n<p class=\"gde-text\"><a href=\"http:\/\/ssrlab.by\/wp-content\/uploads\/2014\/12\/eng_All_About_Natural_Language_Processing.pdf\" class=\"gde-link\">Download (PDF, 285KB)<\/p>","protected":false},"excerpt":{"rendered":"<p>Natural Language Processing is the machine handling of written and spoken human communications. Methods draw on linguistics and statistics, coupled with machine learning, to model language in the service of automation.<\/p>\n<a class = \"excerpt\" href=\"https:\/\/ssrlab.by\/en\/2708\">Read more...<\/a>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[323],"tags":[],"_links":{"self":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts\/2708"}],"collection":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/comments?post=2708"}],"version-history":[{"count":71,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts\/2708\/revisions"}],"predecessor-version":[{"id":3961,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/posts\/2708\/revisions\/3961"}],"wp:attachment":[{"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/media?parent=2708"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/categories?post=2708"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/ssrlab.by\/en\/wp-json\/wp\/v2\/tags?post=2708"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}