The Web 3.0's Pulse : Semantic Web Trends

Currently Hot: Facebook OpenGraph Protocol

Saturday, August 14, 2010

Jena Tutorials

The HP's Jena Framework has become very popular lately in the Semantic Web development world . I started to notice that many of this blog's visitors came here looking for sample code about it. This is why I thought that I could help them out by providing few links where tutorials can be found.

I have noticed that many of the visitors come here seeking some tutorial or getting-started code samples about Jena , besides the code placed in its documentation. Therefore, I decided to look around the Internet and see if I can find some people who already spent some time writing tutorials about Jena . Fortunately, I found some code, which, in my opinion, could very useful if you happen to be someone starting off writing Semantic Web applications. So, here is a small list of links where Jena tutorials can be found:
I agree though, that some of those tutorials may be a few years old, but, in fact, they might serve you well in seeing some kick-off code where the Jena documentation lacks it. For example, in one of those tutorials, I saw a  working set of Jena rules - whereas its documentation was not rich on code about it. 
In case you are wondering what Jena is (which I doubt), I wrote an article describing this Semantic Web Framework a couple of months ago. 

Dear readers, have you ever worked with Jena? What is your personal experience about this framework? Would you favor some other Semantic Web framework instead? Why or why not?

Friday, August 6, 2010

Facebook Questions - Another Search Frontline

Facebook is testing the new Questions  feature - a search engine that finds relevant people to answer other people's questions. Interesting, the word "relevant" here plays great role - Facebook must be pulling data and make conclusions based on that data. "Where from ?" - One may ask. Well, the countless "like"s people do, provide that data, which is now semantically annotated. But behind the scenes, this application concerns several rivals, including Google and Twitter.

The People Search Engine

There are trials to revolutionize the way people search the today. In a world, dominated by authoritative, fast and amazingly complex algorithm for textual search, it seems there is not much to be improved. And for now, people are happy - they just type in what they are interested in, and Google (by saying Google, I also count Yahoo!, Bing and others in)  and: BANG! - it appears in the first 4 results displayed. Although today's search engines subtly show their power by answering well even if the user misspells the word or moreover, categorizing it like Posts, Tweets, Images etc. Google nowadays is even faster in updating its search index, reaching the Real Time Web Informer status.


But however, there are situations in which textual searching can't do much help. Queries like: "Hey Google, what are the songs that have reached No. 1 in UK's Chart in the last 10 years ?" or ... "Bing buddy, I am looking for a movie tonight. Basically I want comedies, like 'American Pie' or 'Dumb and Dumber', do you happen to have some recommendations for me?" etc. etc. (Examples for such situations can be quite a few, I will not go any further). Well, some of these queries might have to wait for the Semantic Search Engine to be built, but the interesting thing is there is a new trend : to build Social (People) Engines, which will not scan text, but people and their habits. These People Engines will try to discover the right person to answer human-only interpretable questions like those above. But to achieve that, these search engines need to have additional metadata about each person : her skills, interests, friends etc. Having that in mind, one can easily conclude that  building such network can be a challenging and expensive task to do - unless... unless it is built by itself - like in the example of Facebook. Facebook is lucky for having so connected network with correctly filled information such as people's names, age, photos, interests, friends etc. In my opinion, Facebook is now preparing to take advantage of that metadata it has: to build internal network for answering questions - people will answer each other's questions on any topic - Facebook will only be the platform to find the "relevant" people to answer them.


Facebook Questions Application


This application is exactly that: means for people to use "Social Search" instead of "Textual Search". A smart move, I say, because engaging real  people in answering topic-specific questions (for free!) is the currently best way to get around the technological gap that prevents us from building software agents to answer the questions for us. Whoever came with the idea of building this application, must have studied people behavior and conclude that people would react on such questions, if they feel they are concerned on some point with the topic of the question - be it their profession or simply a good or bad experience of some product. We will still need to wait to see how this invention of Facebook will impact the Tweetosphere and the ordinary search.

Google's Social Search Engine Efforts


Surprise, surprise, but Facebook did not invent this whole People-answering-questions thing. Google has been spending time on this field quite a while, resulting in an experimental application in Google Labs, which returns 20 % of the results from your Google queries as answers based on your social graph within Google and in inquiry of Aardvark (the closest relative of Facebook Questions app). Aardvark gathers each person's interests they type in and parses the question text, matches entities with people's interests and finds people that might be able to answer the questions. I have been looking into this application for a while and indeed it has proven itself to be very useful: most of the times it did find a person to be able to answer my question and ... what I really like about is that it integrates with the IMs : be it MSN Messenger, GTalk or Skype ... Pretty cool.
Therefore, I think that these kind of engines do have bright future, no matter what vendor creates them.

Readers, what do you think ? Would you use some of these services as your secondary search engines ? Do you think they will one day integrate with the textual search engines ?

Wednesday, July 21, 2010

Facebook Open Graph and Semantic Web Presentation

Here is a thorough and illustrative presentation about  Facebook's Open Graph and it's possible future directions. Presentation's author Matteo Brunati comes to similar conclusions about building an advertising network around the social graph. Have a look by yourself.



Monday, July 19, 2010

Google Acquires Metaweb (Freebase)

Google decides to act quickly and openly step into semantic technology - it acquires Metaweb , the owner of  Freebase , the open knowledgebase. Having this in consideration, it is evident that they are striking back Facebook and their OpenGraph API .It seems like the clash of the web titans has begun...

The Empire Strikes Back

Google has been ignoring Semantic Web long enough. They probably didn't take competition seriously and they didn't feel like their throne on the Web is endangered. But Google could not ignore Semantics any more. Algorithms must be changed. The system evolves, and if they do not want to perform the change, others gladly will. Shortly after Facebook presented their OpenGraph API and their new social bookmarking approach which allows them to build semantically annotated profiles of each of us, Google decides to grab Freebase under its hood. I have been studying Freebase before and according to my opinion, it has the potential of being one of the few data hubs that in future will serve for interconnection of heterogeneous systems. For example, if  two systems refer to same objects, but they describe them differently, by connecting their entities to a single entity on Freebase, the Web will be aware that those two entities are in fact the same, despite their different description. Freebase and DbPedia promise to be the new convention of describing resources in the Web. Googlers are smart, now they will play major role within linking data around the Web - ensuring their importance in this new field.




But will Google utilize this acquiring for search enhancements ? Will they make substantial changes to their keyword search-based tool ? Do Googlers feel threatened ? We are not sure yet.

One thing is for certain. Since large companies start to adopt Semantic Web technologies... it is evident that soon it will become the mainstream. The Semantic Web has waited long enough...

Thursday, July 8, 2010

Facebook Tagging - Semantic Photo Annotation for Free

What is behind the photo tagging feature on Facebook? This seemingly meaningless feature, now allows Facebook to gain digital knowledge of how each person looks, even if the user herself hasn't uploaded a photo of herself yet. But what this feature also means, is that ordinary people around the world, spend their free time to uploading photos, tagging their friends in various positions and occassions. Do you know what this means ? That resources (photos) on Facebook are being annotated manually with high accuracy for free! It seems like we are not far from our digital avatars, after all...

Hidden Semantics in Photos

Many people love the Facebook photo tagging feature : the user basically uploads a photo, tags her friends on the photo and they get pleasant notification that they have been tagged on a certain photo and everyone is happy. But what most people are not aware of, is that by adding tags to photos, they provide Facbook with ground to apply facial recognition algorithms and even more: the hard and manual job of annotating resources with their semantic meaning with at least 95% proven accuracy! Many cannot even recognize it: every time you tag someone on a photo, you basically tell Facebook, the following information:

On the resource [X], the resource [Y] is located on the resource [Z] , 
where [X] is of type Photo, [Y] is of type Person , [Z] is of type Coordinate.
[Y] and [Me] are friends. 
OPTIONAL: [X] is published on the resource [P] , where [P] is a page,
which is in turn relevant to resources [P1,P2...Pn] where [P1] is [SportsTeam].
[Y] likes the resource [P].

Pretty much metadata in a simple tag isn't it ? A semantic software agent (with existent technology) can easily deduce that there is a good chance that you (or the resource [Me]) likes the sports team P1, so why shouldn't it try to suggest you the page P1 ? What if there are other people on the photo that you are not friends with ? Then it is certainly wise to suggest them as your friends since you have been tagged on N photos together, but you are still not friends. This is only a sample of the usage of the semantics embedded within the photos. This system is unique from all aspects. Besides the fact that is free, it is also self -cleaning! What if someone faulty tags to some people? Those people will disagree and remove those tags. Very impressive mechanism, and do not forget it is part of the OpenGraph!

Who still needs to draw photo robots ?

The new exciting new Facebook feature of aiding users while they are trying to tag their friends set off a lot of dust lately. Some users are scared as they now are aware that Facebook recognizes the people in the photo, and will easily gain knowledge on who is on the photo.Imagine how easy is gathering sample set photos for constructing digital profile of each of us:

Give me the photos where [X] is tagged (optional: in the center  of the photo ) :)

Then facial recognition algorithms can be applied and ... ta daam... your face is recognized, along with your real name, surname, who your friends are,  what you like .. interetsting info, don't you think ? Another interesting thing: let's say that you refuse to upload photo of yourself on Facebook. Do you think you can avoid face recognition ? Unfortunately, no, you can't. Your friends will once in a while upload a photo of you and tag you there. So in a good percentage of the photos it will be you and they could draw an image of you even without you uploading a single image of yourself. Oh, and I forgot, a photo where you are tagged is placed as your personal photo by default :) . Creepy feeling  Way to go, Facebook


What do you readers think about Facebook tagging ? Does it really bring extra information to the "Big Face"? Is one's privacy violated in such case ?

Tuesday, July 6, 2010

Google, meet Facebook's OpenGraph Search

With the introduction of the OpenGraph Protocol, Facebook introduces a new concept of searching throughout the Web. Facebook's search works based on what they call "connections" between resources from their OG ontology. Their algorithms are capable of discovering related items to the input query, based on the individual's social neighborhood or perhaps on the frequency of hitting the (now famous) "Like" button. Sounds like a bundle of possibilities, doesn't it ?

FaceRank, the Social Relevance Algorithm

Of course, this is not something Facebook officially announced (it would sound corny, don't you think ?), but the point is, after long 10 years, finally there is a serious candidate to best the PageRank, or at least complement with it. But Google works fine, the whole world searches, people are happy! Why would anyone use Facebook's new lab gadget instead of tested, proven, mature, lightning-fast and precise tool? Well because, there are queries that Google Search simply cannot satisfy! Moreover, their results are based on statistical methods, no people are involved there. What makes Facebook different is the capability to deliver real-time results , fresh and relevant , without deploying complex calculations . If some event is popular, people will rapidly talk about it. Same as with Google, it will be up to the web masters to annotate their web pages with the metadata, but the key differential factor here is that Facebook has the feedback from the users. It can use the number of "Like" hits to give weight to popularity of some particular web page. What if someone puts false metadata? (One of the biggest problems in the Semantic Web, too). In this case, the answer is simple: people will not like it, they will simply ignore it if it is misleading, hence it will be less popular and will have lower positioning. Another advantage from using this approach is that metadata now contains the context of the resource, opening the gates for bringing the conventional Semantic Web Dream . Facebook  is now able to interpret user's query, does she search for related books, movies, sport teams, people... you name it, it finds it... in real time. As written in Times: Google, This Time, Its Personal.

Hey Mark, Recommend Me a Movie, Please

When someone says: "Yeah, the idea of the Semantic Web is great, but if it so wonderful, how come there are no applications to massively leverage it? You say the technology is available for a while.", usually made some point, but I think not anymore. With Facebook's ultimate way of Social Bookmarking, it becomes easily calculable of what users could want, on individual level ! How, you may ask ?
Here is what I am at. (This may be a real idea for semantic application, too). Suppose you want to watch a movie, but you are not really sure what you want to watch... Naturally, you would ask your friends or you would search through the Internet a bit to see where is the movie hype cloud at the moment... (did you realize I said, "at the moment"? Hang on.). Now imagine a widget, that simply communicates the Facebook via OpenGraph API, to check what movies do you like. The widget also supposes that since you like those movies, you have probably watched them, so it makes no sense to suggest them to you again. But how difficult it is, to write a query that says:

"Give me the most popular movies that are related to the comedies I like". We define "related to" as a simple rule: "A movie is related to another if X people that watched the first movie also watched the second. The movie gains ranking in relatedness if at least Y of that people are my friends. The movie gains ranking if there are at least Z pages with more than 50 likes on the Web". 

Hmmm, not so difficult to be written in a query language. For now some of these aspects are not covered in the OpenGraph ontology (I refer to the Movie Genre), but undoubtly, it could easily be added. On the other side, for the application user, it is as simple as logging in to Facebook, and pressing the "Recommend" button. Welcome to the Semantic reality, Neo. Btw, how do you write "My favorite movies" in Google ? :)

But appart from the interesting search ideas the OpenGraph brings, my deepest beliefs are that Facebook's reason number one to introduce this protocol has e-Marketing roots i.e. to deliberately interfere with Google's primary business model - with personalized, perfect ad targeting tool .

What do you think ? Will this Facebook API bring new methods of warfare between the web titans ? Will it provide better searching for end-users ? Will ultimately, data find us ? How will Google eventually respond ? Is this the final gate that needed to be opened, for semantic applications to be massively written ?

Sunday, July 4, 2010

Facebook and the Semantic Web: Weaving the Social or the Advertising Graph?

You probably already heard about the Facebook's new OpenGraph Protocol. It represents a new way of making connections between topics people like around the web, thus embedding metadata within the webpages itself. Why is Facebook doing this ? Does it want (really) to act as a social hub platform for bridging the Semantic Web to reality?



Finally a big player enters the Semantic Web realm. One that is recognizable all over the world. One that people have confidence in. One that promises to be powerful enough, to integrate topics from different webpages and connect them to corresponding people. One that will get rid of the chicken-and-the-egg vicious circle of Semantic annotation and Semantic Applications. One graph to rule them all : Facebook's OpenGraph.

Facebook, the Chicken and the Egg

Facebook apparently is trying to motivate webmasters to start embedding semantics into webpages, similarly to how meta keywords and meta description tags are embedded today for SEO. That would eventually give the desired push and stable ground for Semantic Applications to be finally built. People are already familiar with this way of embedding metadata, thus the motivation for them lies in the fact that Facebook will utilize that metadata whenever someone puts the mouse over the link that describes how a person "likes" something. But what is happening in background ? Is this simplified mapping to Facebook's ontology one step driven by the desire for people to share what they really like around different platforms ?  Does Facebook have hidden intentions in this whole story ?

The Impact on Ordinary Users

Well, what do average users get from the Social Graph ? Of course, they could leverage this new feature in order to spread the word about services/products they prefer or offer, providing additional fuel to marketing in Social Media. From that aspect, users will get even more specific recommendations from friends about things that might interest them. Of course, friends have similar interests and there is a good chance that they will at least be intrigued about what one's friends like. Moreover, "like" web sites that aggregate Facebook page titles and groups have begin to emerge. Some users find this aggregation amusing.

Facebook Flaws in Semantics : Why ?

As it was recently published in a post on Read Write Web , Facebook did leave flaws in embedding semantics in web pages. Some of them are known to Semantic Web enthusiasts from long time ago, such as the ambiguity problem when identifying resources. In terms of the OpenGraph protocol, there is no means to denote that two resources on the Web refer to actually the same thing. Therefore, integration between heterogeneous systems is not easy at all. Secondly, items with same names refer to same things although they point to different terms. This means there is no way to denote that a page is relevant to the car Jaguar, not the animal jaguar. Furthermore, the OpenGraph leaves no way to build relations between resources, assuming that the only relation is : is_relevant_to . This relation applies to web pages and items and items to people, respectively. This conclusion comes since there is no way to embed multiple objects into a single web page.

The Open Advertising Protocol

This is not something that Facebook publicly says, but if one gets into little deeper thinking, becomes obvious. Facebook is not concerned about allowing people brag to the others what they like. The company is concerned about mapping the users' interests in another graph, which I take the freedom to name it Open Advertising Protocol. It refers to a graph that will try to make connections between topics that might interest the user and her social graph, individually and in groups. What this means is the following: Facebook is trying to gain information about the meaning of the things because it needs more precise targeting for its personalized ads! It is fairly simple. Every time a user presses the "Like" button, Facebook gains insight on that user's interests! By having this knowledgebase at hand, Facebook will soon have enough data to improve their Ad targeting algorithm. What is even scarier, even if one does not press the "Like" button, they will be able to map your interests roughly based on your friend's interests! One might think: Fine, then people will eventually stop hitting that button once they realize this. But hold on a second! Facebook was created to fulfill a human need for social interaction, an interaction that was not satisfied by any other media before ! The point is, people are not that inert as one might think! "Like" it or not, the Big Face will be able to find out who we are, who we hang out with, what are we interested in aaaand ... what companies have better chance of selling to us!

What do you think? What is the reason for Facebook to enter this Semantic Web Game ? Why does it leave flaws, although it has both knowledge and infrastructure to make it differently? Does it really want help people share things they like or this is just a preparation for the Perfect Advertising Tool and even bigger profit ?

Friday, October 30, 2009

Semantic Web and Peer-to-Peer Networks

Decentralized Knowledge Management and Information Exchange

The idea for the Semantic Web  brings interesting ideas for revolutionary approaches in the fields such as knowledge management. Knowledge Management (KM) becomes one of the essential driving forces for the existence of large communities. However, searching through the knowledge bases is very limited. The Semantic Web technologies promise concept alignment and powerful search features to be incorporated into knowledge management systems. But due to the client-server architecture upon which KM systems most often rely on, appearance of physical bottlenecks are very likely to happen. Therefore, an alternative architecture appropriate for large-scale systems is needed. As you already may guess, here flat networks come into play - the Peer-to-Peer (P2P) systems. 

P2P systems are not only scalable, but improve their performances as new nodes join the network. This unique feature gives them suitable ground for all sharing software applications. However, searching through these networks is limited to keyword search only, and different peers may name same resources differently, so it is obvious that it lacks semantics when resources with anonymous peers.

By combining Semantic Web technologies and P2P networks, could possibly bring ultimate conditions for advanced knowledge management techniques. Distributing knowledge accross different transparent physical locations and also providing peers to use their own views upon same data, can bring tremendous change in the evolution of the Semantic Web.  At first glance, it seems like an extraordinary idea - it would be wonderful if I could share a file with other peers who would be able to understand what it is and how it is related with other files someone else owns. Also, KMSs could distribute their knowledge on different peers, possibly with previous domenization of the knowledge in order to boost performances.

But if peers are free to annotate the resources they share with others, how will the knowledge ( ontology ) alignment occur ? There must be some kind of common mechanism for annotation so that alignment happens automatically and peer agents will be able to identify the relation of the resources shared. One suggested approach is to use small common vocabularies among peers. Then the peer agents will be able to align the knowledge automatically. But the question of performance for such large-scale system is yet to be proved as no real Semantic P2P application has been deployed.

Note: For anyone interested in this topic, I recommend the book: Semantic Web and Peer-to-Peer: Decentralized Management and Exchange of Knowledge and Information. I will continue the discussion once I finish reading this book.

Sunday, October 25, 2009

Jena, a Framework for developing Semantic Web Applications


Jena, Semantic Web framework, advantages and features

 Jena is a Java framework for developing Semantic Web applications. It has been developed by HP Labs and it is an open source project. Basically, Jena provides Java environment for working with RDF, RDFS, OWL, SPARQL and reasoning engines. The Jena framework creates an additional layer of abstraction that translates the statements and constructs of the Semantic Web into Java artifacts, such as classes, objects, methods and attributes. These artifacts reduce the effort needed for programming Semantic Web applications. One of the strongest sides of Jena lies in its excellent documentation. The exhaustive resources, including descriptions and tutorials that can be found on the Web encourage programmers to further develop their Semantic Web applications utilizing this framework.
As part of its RDF features, Jena offers managing with RDF resources, writing them in RDF/XML, N3 and N-Triples format. Jena also supports working with the RDF Schema, by providing API for all the vocabulary extensions it brings. Moreover, Jena covers the usage of OWL, in one of the three variants: Full, Description Logic, and Lite. The OWL API provides the ability to navigate through the graph, locate resources and retrieve them from the model. Regardless to the schema and the data models (which can be separate resources) used, Jena can simultaneously work with multiple ontologies from different sources. The API which comes with the framework makes the knowledge sharing process extremely easy, as every resource comes with its URI, Jena is excellent in working with the knowledge shared across the (Semantic) Web. The framework also covers methods for validating an ontology and derivation logging, which enables the developer to see how Jena concludes the answers of the query.
Regarding the persistence storage, Jena perfectly works with files containing OWL or RDF data, but has an API for database backend as well. Because of its high level of generics crafted into its software design, Jena can be easily bound to SQL databases from different vendors. All a developer needs is an appropriate driver for the particular SQL database.
Querying the knowledge graph is an important topic when discussing semantic web frameworks. Jena supports querying the model through the API, or by directly constructing SPARQL query to retrieve the results. The knowledge base can be attached to a web server designed especially for Jena, named Joseki (www.joseki.org). Joseki acts as a mediator between the SPARQL query input through GET or POST HTTP methods, and returns RDF/XML response with the results, which can be further formatted with XSLT.  
Perhaps the most powerful component of the Jena Framework is the Inference API. This API contains several reasoner types, which efficiently conclude new relations in the knowledge graph. Among the reasoners, there are: RDF(S), OWL, Transitive and Generic reasoners. It is worth mentioning that Jena is compatible with third party reasoners, such as the Pellet reasoner. All of the reasoners can be configured individually, by creating special resources that contain the desired configuration and  then using it to perform the reasoning. For example, the reasoned can be configured to run in forward-chaining or backward-chaining mode, or an OWL reasoner can be instructed to use a Description Logic (or Full or Lite) memory model specifically in the favor better reasoning performance.

Disadvantages

Despite the powerful abilities and the high level of abstraction provided, Jena has some serious disadvantages. For instance, when retrieving datasets, the framework places all statements into the main memory, often causing an overflow in the heap of the Java Virtual Machine (JVM). Therefore, the needs significant amount of space, depending on the number of statements that are retrieved in the resulting data set. This is also true even if one decides to use SQL database for persistence.
The second disadvantage is regarding the threading. Namely, Jena is not thread safe and consistency and concurrency issues can easily occur. The API provides methods for declaring critical regions but it is up to the programmer to take care of the threads using the model.
The third, and possibly the most relevant disadvantage is the cost of the inference process. Inference capability is one of the basic features of the knowledgebase and yet the most powerful one. Without inference, a knowledgebase would not be much different from an ordinary database. As mentioned earlier, the reasoning process infers implicit statements in the knowledge graph. Hence the number of edges in the graph rapidly increases, requiring more time to navigate and locate a specific resource from it. Adding large number of statements in the knowledge model is a time- and memory-consuming process. However, efforts are being made to decrease these high costs by using methods known as graph closure and graph reduction.

Summary

The Jena Framework is an excellent tool for managing resources needed for the Semantic Web applications. Being developed in Java, it is applicable to various environments. In addition, it is open source and strongly backed up by solid documentation. Even though it has some significant disadvantages, it is still one of the most powerful frameworks for Semantic Web technologies and holds the potential to become de facto standard when it comes to developing such programs. Frameworks like Jena are worth investing in, since they might play the key role in the evolution of the WWW into Semantic Web, predicted by Sir Tim Berners Lee.

Saturday, October 24, 2009

Alchemy from Raw Text to Semantic Annotations

Alchemy API, a free web API that can be used to semantically annotate web resources such as HTML documents or pure raw texts. The  Alchemy API, uses NLP (Natural Language Processing) to extract the meaning of the input text. Alchemy is capable of retrieving entities, keywords, pure text txtraction, text categorization, language identification and probably some other features that I haven't explored yet. The response of the API is ordinary XML ( RDF included here as well ), which of course, is easy to parse and integrate with any platform. Unlike OpenCalais, which offers typical SOAP web service( which is also a tehcnology I personally admire, you can read my post about it), Orchestr8's Alchemy API uses web requests to their server to invoke their NLP technology. The API is well documented, with downloadable examples for a variety of technologies. Since I am ASP.NET developer, I only downloaded the C#.NET SDK examples and it seems really neat and simple to use. It is very probable that other SDK's are as good as the mentioned one. Please note that you would need  an API key before using this service. This Orchestr8's API, in my opinion bears a lot of potential in dramatically decreasing the effort for semantic annotation of world's web content, which is most likely the stepping stone for the real semantic evolution. An interesting thing that I noticed in the Alchemy's response is that actually contains some kind of entity alignment in it. For instance, it is capable of disambiguation of a given term, so  in the RDF response there is an OWL ontology alignment piece( powered by LinkedData I suppose ), that aligns the disambiguated entity with semantic resources from sources like DBPedia, Freebase, CIA Factbook,GeoNames etc. Really cool. That could boost semantic knowledge management in even more advanced semantic applications.
Orchestr8 offers the API in several packages. The basic package is free and allows 30 000 calls per day, which, is enough for small to medium applications. If a good idea for a semantic web application is born, this certainly will not be the limit.



Thursday, October 15, 2009

Triplify: Expose Your Information as RDF

Yet another way into automating the process of semantic annotation of your web content. As many times mentioned, one of the biggest obstacles for the Semantic Web to become fully implemented, is that people (web masters) are forced to manually annotate every resource they have on the Web in order (future) applications to use that information. The trouble is the term future applications. There is no killer semantic web application yet, partly because there is no semantic data available through the Web.
So here comes Triplify into play. Triplify is a framework that works over your SQL database and the web master selects the columns that are of interest when constructing the triples, and it automatically generates RDF using the pattern specified by the owner.




This is an illustration on how Triplify  works:



Disadvantages
Besides the nice approach of generating RDF triples, some serious shortcomings are hidden under the hood. 
First, it has performance issues and is aimed for using on small to medium web sites. 
Secondly, it is still implemented only in PHP. Semantic Web goes far beyond technology limitations. However, more developers are needed for implementations on other platforms.

Conclusion 
 As far I explored, it seems fairly easy to integrate, and it might be very helpful for exporting your data in RDF. But one question still remains.
What to do with RDF ? How does one utilize RDF to make semantic killer application?

Monday, October 12, 2009

The "Prophet" Himself: Tim Berners Lee Talks About the Semantic Web

Here is an interesting video for all those of you who would like to see Tim Berners Lee talking about the phenomenon he predicted about 10 years ago and the revolution which we all anxiously wait to be triggered. I think that much of what he says in the video most of you have already read or heard about, but it is still a nice feeling to see him talking about the Semantic Web. He is so passionate talking about the Web of Data, he really believes.




Sunday, October 11, 2009

Meet Sindice - the Semantic Web Index

Sindice is a Semantic Web index with a search engine. . (Here is the link: sindice.com). Interestingly, one of its authors is Nova Spivack from Radar Technologies, the same guy who runs twine.com. Is this the semantic search engine twine talks about ? Sindice claims to offer semantic search for terms and properties or triples, but personally, I can hardly notice the benefit of using it - It is only capable of indexing things, it is not capable of extracting terms' relations with other terms - something the Semantic Web is about. Why would I need a semantic index or a search through that index. Probably the best application of these services would be making of semantic piped tasks. The Sigma search engine is aggregating definitions for the terms from different sources, and what I really like about it is the ability for the users to control the sources of the definitions. Sigma can export the search results in various formats, such as RDF or JSON, but I am really having hard time to see the benefits of it without getting the relations of the terms.


In general, Sindice  has some  fancy marketing, but what it really has under the hood remains to be seen. Personally I don't think there is much to brag about.

Insights of the Semantic World: APIs, Freebase and Collaborative Semantic Web

The Semantic Web is all about collaboration. Collaboration between machines, collaboration between people, and collaboration between machines and people. Everything (r)evolves around this term. Though this blog does not intend to provide tutorial-like information about the Semantic Web, this presentation I came across on slideshare really drew my attention, so I decided to share it here. Very nice way of organizing the concepts about the Semantic Web and oh, the last slide is simply amazing :) . Not just because of the point where the arrow points to, but because of the curve we are about to climb on. When I see slideshows like this, I always get excited and say to myself: "Wow! What a brilliant idea! This is ... great! You know, this could change the world! It's such a wonderful feeling!". My suggestion is that you take a look at the presentation and see for yourself. 


Friday, October 9, 2009

What Would You Do With RDF Organized Knowledge ?

So, lets say that there is an ideal tool for creating RDF from your HTML pages. And now what ? This question came to my mind... What is the first application you would write knowing that many sites publish their knowledge in RDF / OWL format ? The sad thing is, I could not answer immediately. So think about it, what would be the amazing benefit of publishing RDF ?
If anyone has implemented such applications, feel free to share them here.

Open Calais: Automatic RDF Annotation of Raw Text

OpenCalais: Automatic Knowledge Extraction

Have you tried OpenCalais? It's a web service that automatically annotates raw text with semantic meaning - it generates an RDF file from it. Well, it's still far from perfect, but it obviously has good performance in semantic annotation of your text. Basically, one can copy & paste the text from her blog and get RDF-structured knowledge based on the text. From OpenCalais say they use NLP(Natural Language Processing) techniques to analyze the text and calculate the relevance of the recognized concepts in the text. It can be very useful for web masters, since this approach can be a true time-saver and one of the easiest steps towards automatic knowledge publication. In my opinion, automating the process of semantic annotation of the web documents from one side and presenting the benefits of the Semantic Web to the web masters are the two crucial steps that need to be taken in order to come closer to the true implementation of the Semantic Web itself. These two steps can end the vicious chicken-and-the-egg circle: Webmasters refuse to put extra effort in embedding knowledge into their web pages, since there are no semantic applications that would use that knowledge and make web masters' life easier. But because there is no semantic knowledge, no real semantic applications can be developed. And this circle goes on and on.
With OpenCalais, you can publish your knowledge via API, so new custom applications can arise from your website, blog, wiki, e-commerce page or similar. I admit I still need to read and play with the Calais to explore its full features, but from what I have seen so far it looks excellent. This article does not aim to advertise OpenCalais in any way, nor I am related to it, but I would like to emphasize the importance of its existence as a service, that could be the stepping stone towards unleashing the power of the Semantic Web.
Moreover, OpenCalais has plugins for Wordpress (ohhh, none of them for Blogger :-( ), to automatically generate tags ( Tagaroo ). OpenCalais also can be integrated with Drupal. Seems like a nice application.
I believe that by using this service, the number of semantically annotated pages will rapidly rise. That will make a good ground for development of even more advanced Semantic Applications. Try the video and go to the site, so tell me what do you think.



Here is how applications can be build on top of it:



You can try the OpenCalais Document Viewer, to see how it generates the RDF output.

It only remains to see if OpenCalais will fulfill its glorious mission. I really recommend OpenCalais to the Semantic Web Community, its effort deserves attention


Monday, September 21, 2009

Thompson-Reuters Claims Is Able to Extract Semantics From Free Text and Export it to Oracle Database

According to the latest news, Thompson-Reuters has reported that with OpenCalais, a metatagging service, will be integrated with an Oracle database. Here is how it works: first a number of raw text (unstructured) documents are identified in a database, filesystem or across a network, then OpenCalais is invoked via a web-service, which returns a set of RDF triples which are back then saved in a RDF triple store.
This probably means a beginning of the end of the extra effort needed to semantically annotate the enormous number of web documents that are deployed all over the Internet. With such possibilities at service, web masters could finally tag their web pages through a single click - and bother no more. Businesses will benefit from this too. Their scattered knowledge bases can now be easily integrated into a single entity - which could possess its own inference engine and further utilize the semantics it gets.
Currently OpenCalais claim that they are processing between 3 and 5 million documents per day, and they will soon attract even more developers to use their service.

Will this be the trigger to catalyze the Semantic (r)evolution ?

Saturday, September 19, 2009

The Semantic Search Engine : Dream 3.0 ?

Today I read about the latest try to fulfill the famous Web Dream: The Semantic Search Engine. Wouldn't it be nice to have such a wonderful tool, that can actually understand you ? You can ask it about anything, it is the Global Mind, it crunches data and comprehends the whole Web, the largest knowledge management application humanity has ever built. And the best thing is, it learns and gets smarter with every day... by itself.
Sounds like a quote from a Science Fiction book, but is it that far ? It's been about 10 years since the publication of the famous paper in Scientific American by Tim Berners Lee, but yet no (r)evolution has occured. There is no single killer application, fueled by the Semantic Technologies. But why ?
The whole computer industry lives for roughly 60 years, the Internet era has begun in the 1990s, so a period of 10 years means a lot of time for the Web. That is huge amount of time. We have the standards, we have the tools, we have the frameworks, the knowledge ...
I have read several articles and it seems there is a logical explanation of this phenomenon: it's the humans that are wrong... (again). It's not the problem in making machines undersand what we mean (personally I think it sounds like the most exciting part when telling someone what is the Semantic Web all about: computers will undersand ? Really ? Like in the movies ? Will I be able to ask them via voice control ? ). The trouble is that people are lazy. The WWW is the biggest and the fastest growing entity on the whole planet. It is enormous. People will need extra effort to annotate all that data across the web. But it is tedious and time-consuming (Hey, didn't we invent computers because of that ?). But it's they that don't understand, not the machines. Machines are ready to learn. Another issue for that would come from the fact that humans are spoiled and selfish - people lie. Yes, they do. There is no rightful force to make webmasters embed true information about their web pages. (Remember the keyword stuffing problem ? ). How will someone even make them want to start annotating the pages ? I believe that here lie most of the problems for the stagnation of the Semantic Web and its applications.

There are efforts to automate the process through Natural Language Processing(NLP) but I wonder if it ever reaches the desired level of automation. Here is a good article about what Oracle does : Oracle & OpenCalais - Semantic Database. This thing really makes me happy because of the burst of hope that the Semantic Web is not an e-Myth.
Back to the search engines. The team of Twine.com has been busy trying to achieve the unimaginable: produce a true Semantic Search Engine. Here is the original post I found T2 - Twine's Semantic Search Engine. If this becomes true, all the hype will disappear in the mist. I understand why people are sceptical, but presonally, guys, you don't know what might happen. Maybe it is possible. Requirements are high - a volatile system, evolving every second, reasoning and comprehending, accurate, fast, robust ... but there is still a chance. The Semantic Engine is the one of the most desired applications of the Semantic Web. It will be a major breakthrough - although many find it tough to believe. Will the openess and sharing prevail at the end ?