The Web 3.0's Pulse : Semantic Web Trends

Currently Hot: Facebook OpenGraph Protocol

Showing posts with label Natural Language Processing. Show all posts
Showing posts with label Natural Language Processing. Show all posts

Saturday, October 24, 2009

Alchemy from Raw Text to Semantic Annotations

Alchemy API, a free web API that can be used to semantically annotate web resources such as HTML documents or pure raw texts. The  Alchemy API, uses NLP (Natural Language Processing) to extract the meaning of the input text. Alchemy is capable of retrieving entities, keywords, pure text txtraction, text categorization, language identification and probably some other features that I haven't explored yet. The response of the API is ordinary XML ( RDF included here as well ), which of course, is easy to parse and integrate with any platform. Unlike OpenCalais, which offers typical SOAP web service( which is also a tehcnology I personally admire, you can read my post about it), Orchestr8's Alchemy API uses web requests to their server to invoke their NLP technology. The API is well documented, with downloadable examples for a variety of technologies. Since I am ASP.NET developer, I only downloaded the C#.NET SDK examples and it seems really neat and simple to use. It is very probable that other SDK's are as good as the mentioned one. Please note that you would need  an API key before using this service. This Orchestr8's API, in my opinion bears a lot of potential in dramatically decreasing the effort for semantic annotation of world's web content, which is most likely the stepping stone for the real semantic evolution. An interesting thing that I noticed in the Alchemy's response is that actually contains some kind of entity alignment in it. For instance, it is capable of disambiguation of a given term, so  in the RDF response there is an OWL ontology alignment piece( powered by LinkedData I suppose ), that aligns the disambiguated entity with semantic resources from sources like DBPedia, Freebase, CIA Factbook,GeoNames etc. Really cool. That could boost semantic knowledge management in even more advanced semantic applications.
Orchestr8 offers the API in several packages. The basic package is free and allows 30 000 calls per day, which, is enough for small to medium applications. If a good idea for a semantic web application is born, this certainly will not be the limit.



Friday, October 9, 2009

Open Calais: Automatic RDF Annotation of Raw Text

OpenCalais: Automatic Knowledge Extraction

Have you tried OpenCalais? It's a web service that automatically annotates raw text with semantic meaning - it generates an RDF file from it. Well, it's still far from perfect, but it obviously has good performance in semantic annotation of your text. Basically, one can copy & paste the text from her blog and get RDF-structured knowledge based on the text. From OpenCalais say they use NLP(Natural Language Processing) techniques to analyze the text and calculate the relevance of the recognized concepts in the text. It can be very useful for web masters, since this approach can be a true time-saver and one of the easiest steps towards automatic knowledge publication. In my opinion, automating the process of semantic annotation of the web documents from one side and presenting the benefits of the Semantic Web to the web masters are the two crucial steps that need to be taken in order to come closer to the true implementation of the Semantic Web itself. These two steps can end the vicious chicken-and-the-egg circle: Webmasters refuse to put extra effort in embedding knowledge into their web pages, since there are no semantic applications that would use that knowledge and make web masters' life easier. But because there is no semantic knowledge, no real semantic applications can be developed. And this circle goes on and on.
With OpenCalais, you can publish your knowledge via API, so new custom applications can arise from your website, blog, wiki, e-commerce page or similar. I admit I still need to read and play with the Calais to explore its full features, but from what I have seen so far it looks excellent. This article does not aim to advertise OpenCalais in any way, nor I am related to it, but I would like to emphasize the importance of its existence as a service, that could be the stepping stone towards unleashing the power of the Semantic Web.
Moreover, OpenCalais has plugins for Wordpress (ohhh, none of them for Blogger :-( ), to automatically generate tags ( Tagaroo ). OpenCalais also can be integrated with Drupal. Seems like a nice application.
I believe that by using this service, the number of semantically annotated pages will rapidly rise. That will make a good ground for development of even more advanced Semantic Applications. Try the video and go to the site, so tell me what do you think.



Here is how applications can be build on top of it:



You can try the OpenCalais Document Viewer, to see how it generates the RDF output.

It only remains to see if OpenCalais will fulfill its glorious mission. I really recommend OpenCalais to the Semantic Web Community, its effort deserves attention


Monday, September 21, 2009

Thompson-Reuters Claims Is Able to Extract Semantics From Free Text and Export it to Oracle Database

According to the latest news, Thompson-Reuters has reported that with OpenCalais, a metatagging service, will be integrated with an Oracle database. Here is how it works: first a number of raw text (unstructured) documents are identified in a database, filesystem or across a network, then OpenCalais is invoked via a web-service, which returns a set of RDF triples which are back then saved in a RDF triple store.
This probably means a beginning of the end of the extra effort needed to semantically annotate the enormous number of web documents that are deployed all over the Internet. With such possibilities at service, web masters could finally tag their web pages through a single click - and bother no more. Businesses will benefit from this too. Their scattered knowledge bases can now be easily integrated into a single entity - which could possess its own inference engine and further utilize the semantics it gets.
Currently OpenCalais claim that they are processing between 3 and 5 million documents per day, and they will soon attract even more developers to use their service.

Will this be the trigger to catalyze the Semantic (r)evolution ?

Saturday, September 19, 2009

The Semantic Search Engine : Dream 3.0 ?

Today I read about the latest try to fulfill the famous Web Dream: The Semantic Search Engine. Wouldn't it be nice to have such a wonderful tool, that can actually understand you ? You can ask it about anything, it is the Global Mind, it crunches data and comprehends the whole Web, the largest knowledge management application humanity has ever built. And the best thing is, it learns and gets smarter with every day... by itself.
Sounds like a quote from a Science Fiction book, but is it that far ? It's been about 10 years since the publication of the famous paper in Scientific American by Tim Berners Lee, but yet no (r)evolution has occured. There is no single killer application, fueled by the Semantic Technologies. But why ?
The whole computer industry lives for roughly 60 years, the Internet era has begun in the 1990s, so a period of 10 years means a lot of time for the Web. That is huge amount of time. We have the standards, we have the tools, we have the frameworks, the knowledge ...
I have read several articles and it seems there is a logical explanation of this phenomenon: it's the humans that are wrong... (again). It's not the problem in making machines undersand what we mean (personally I think it sounds like the most exciting part when telling someone what is the Semantic Web all about: computers will undersand ? Really ? Like in the movies ? Will I be able to ask them via voice control ? ). The trouble is that people are lazy. The WWW is the biggest and the fastest growing entity on the whole planet. It is enormous. People will need extra effort to annotate all that data across the web. But it is tedious and time-consuming (Hey, didn't we invent computers because of that ?). But it's they that don't understand, not the machines. Machines are ready to learn. Another issue for that would come from the fact that humans are spoiled and selfish - people lie. Yes, they do. There is no rightful force to make webmasters embed true information about their web pages. (Remember the keyword stuffing problem ? ). How will someone even make them want to start annotating the pages ? I believe that here lie most of the problems for the stagnation of the Semantic Web and its applications.

There are efforts to automate the process through Natural Language Processing(NLP) but I wonder if it ever reaches the desired level of automation. Here is a good article about what Oracle does : Oracle & OpenCalais - Semantic Database. This thing really makes me happy because of the burst of hope that the Semantic Web is not an e-Myth.
Back to the search engines. The team of Twine.com has been busy trying to achieve the unimaginable: produce a true Semantic Search Engine. Here is the original post I found T2 - Twine's Semantic Search Engine. If this becomes true, all the hype will disappear in the mist. I understand why people are sceptical, but presonally, guys, you don't know what might happen. Maybe it is possible. Requirements are high - a volatile system, evolving every second, reasoning and comprehending, accurate, fast, robust ... but there is still a chance. The Semantic Engine is the one of the most desired applications of the Semantic Web. It will be a major breakthrough - although many find it tough to believe. Will the openess and sharing prevail at the end ?