Friday, May 22, 2009

Overhauling the Java UTF-8 charset

he UTF-8 charset implementation, which is available in all JDK/JRE releases from Sun, has been updated recently to reject non-shortest-form UTF-8 byte sequences. This is because the old implementation might be leveraged in security attacks. Since then I have been asked many times about what this "non-shortest-form" issue is and what the possible impact might be, so here are some answers.

The first question usually goes: "What is the non-shortest-form issue"?

The detailed and official answer is at Unicode Corrigendum #1: UTF-8 Shortest Form. Simply put, the problem is that Unicode characters can be represented in more than one way (form) in the "UTF-8 encoding" than many people think or believe. When asked what UTF-8 encoding looks like, the simplest explanation would be the following bit pattern:

# Bits Bit pattern
1 7 0xxxxxxx
2 11 110xxxxx 10xxxxxx
3 16 1110xxxx 10xxxxxx 10xxxxxx
4 21 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

The pattern is close, but it's actually wrong, based on the latest definition of UTF-8. The preceding pattern has a loophole in that you can actually have more than one form represent a Unicode character.

For ASCII characters from u+0000 to u+007f, for example, the UTF-8 encoding form maintains transparency for all of them, so they keep their ASCII code values of 0x00..0x7f (in one-byte form) in UTF-8. Based on the preceding pattern, however, these characters can also be represented in 2-bytes form as [c0, 80]..[c1, bf], the "non-shortest-form".

The following code shows all of the non-shortest-2-bytes-form for these ASCII characters, if you run code against the "old" version of the JDK and JRE (Java Runtime Environment).


byte[] bb = new byte[2];
for (int b1 = 0xc0; b1 < 0xc2; b1++) {
for (int b2 = 0x80; b2 < 0xc0; b2++) {
bb[0] = (byte)b1;
bb[1] = (byte)b2;
String cstr = new String(bb, "UTF8");
char c = cstr.toCharArray()[0];
System.out.printf("[%02x, %02x] -> U+%04x [%s]%n",
b1, b2, c & 0xffff, (c>=0x20)?cstr:"ctrl");
}
}

The output would be as follows:


...
[c0, a0] -> U+0020 [ ]
[c0, a1] -> U+0021 [!]
...
[c0, b6] -> U+0036 [6]
[c0, b7] -> U+0037 [7]
[c0, b8] -> U+0038 [8]
[c0, b9] -> U+0039 [9]
...
[c1, 80] -> U+0040 [@]
[c1, 81] -> U+0041 [A]
[c1, 82] -> U+0042 [B]
[c1, 83] -> U+0043 [C]
[c1, 84] -> U+0044 [D]
...

So, for a string like "ABC", you would have two forms of UTF-8 sequences:


"0x41 0x42 0x43" and "0xc1 0x81 0xc1 0x82 0xc1 0x83"

The Unicode Corrigendum #1: UTF-8 Shortest Form specifies explicitly that "The definition of each UTF specifies the illegal code unit sequences in that UTF. For example, the definition of UTF-8 (D36) specifies that code unit sequences such as [C0, AF] are illegal."

Our old implementation accepts those non-shortest-form (while it never generates them when encoding). The new UTF_8 charset now rejects the non-shortest-form byte sequences for all BMP characters. Only the "legal byte sequences" listed below are accepted.


/* Legal UTF-8 Byte Sequences
*
* # Code Points Bits Bit/Byte pattern
* 1 7 0xxxxxxx
* U+0000..U+007F 00..7F
* 2 11 110xxxxx 10xxxxxx
* U+0080..U+07FF C2..DF 80..BF
* 3 16 1110xxxx 10xxxxxx 10xxxxxx
* U+0800..U+0FFF E0 A0..BF 80..BF
* U+1000..U+FFFF E1..EF 80..BF 80..BF
* 4 21 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
* U+10000..U+3FFFF F0 90..BF 80..BF 80..BF
* U+40000..U+FFFFF F1..F3 80..BF 80..BF 80..BF
* U+100000..U10FFFF F4 80..8F 80..BF 80..BF
*/

The next question usually is: "What would be the issue/problem if we keep using the old version of JDK/JRE?"

First, I'm not a lawyer — oops, I meant I'm not a security expert:-) — so my word does not count. We consulted with our security experts instead. Their conclusion is that while "it is not a security vulnerability in Java SE per se, it may be leveraged to attack systems running software that relies on the UTF-8 charset to reject these non-shortest form of UTF-8 sequences".

A simple scenario that might give you an idea about what the above "may be leveraged to attack..." really means:

  1. A Java application would like to filter the incoming UTF-8 input stream to reject certain key words, for example "ABC".
  2. Instead of decoding the input UTF-8 byte sequences into Java char representation and then filter out the keyword string "ABC" at Java "char" level, for example:

    String utfStr = new String(bytes, "UTF-8");
    if ("ABC".equals(strUTF)) { ... }
    The application might choose to filter the raw UTF-8 byte sequences "0x41 0x42 0x43" (only) directly against the UTF-8 byte input stream and then rely on (assume) the Java UTF-8 charset to reject any other non-shortest-form of the target keyword, if there is any.
  3. The consequence is the non-shortest form input "0xc1 0x81 0xc1 0x82 0xc1 0x83" will penetrate the filter and trigger a possible security vulnerability, if the underlying JDK/JRE runtime is an OLD version.

So the recommendation is: Update to the latest JDK/JRE releases to avoid the potential risk.

Wait, there is another big bonus for updating: performance.

The UTF-8 charset implementation has not been updated or touched for years. UTF-8 encoding is very widely used as the default encoding for XML, and more and more websites use UTF-8 as their page encoding. Given that fact, we have taken the defensive position of "don't change it if it works" during the past years.

So Martin and I decided to take this opportunity to give it a speed boost as well. The following data is from one of my benchmark's run data, which compares the decoding/encoding operations of new implementation and old implementation under -server vm. (This is not an official benchmark: it is provided only to give a rough idea of the performance boost.)

The new implementation is much faster, especially when decoding or encoding single bytes (those ASCIIs). The new decoding and encoding are faster under -client vm as well, but the gap is not as big as in -server vm. I wanted to show you the best:-)


Method Millis Millis(OLD)
Decoding 1b UTF-8 : 1786 12689
Decoding 2b UTF-8 : 21061 30769
Decoding 3b UTF-8 : 23412 44256
Decoding 4b UTF-8 : 30732 35909
Decoding 1b (direct)UTF-8 : 16015 22352
Decoding 2b (direct)UTF-8 : 63813 82686
Decoding 3b (direct)UTF-8 : 89999 111579
Decoding 4b (direct)UTF-8 : 73126 60366
Encoding 1b UTF-8 : 2528 12713
Encoding 2b UTF-8 : 14372 33246
Encoding 3b UTF-8 : 25734 26000
Encoding 4b UTF-8 : 23293 31629
Encoding 1b (direct)UTF-8 : 18776 19883
Encoding 2b (direct)UTF-8 : 50309 59327
Encoding 3b (direct)UTF-8 : 77006 74286
Encoding 4b (direct)UTF-8 : 61626 66517

The new UTF-8 charset implementation has been integrated in JDK7, Open JDK 6, JDK 6 update 11 and later, JDK5.0u17, and 1.4.2_19.

If you are interested in what the change looks like, you can take a peek at the webrev of the new UTF_8.java for OpenJDK7.


Thanks to the original author and the source

Security: Quick tour of controlling Applets

Security

Trail: Security Features in Java SE

In this trail you'll learn how the built-in Java™ security features protect you from malevolent programs. You'll see how to use tools to control access to resources, to generate and to check digital signatures, and to create and to manage keys needed for signature generation and checking. You'll also see how to incorporate cryptography services, such as digital signature generation and checking, into your programs.

The security features provided by the Java Development Kit (JDK™) are intended for a variety of audiences:

  • Users running programs:

    Built-in security functionality protects you from malevolent programs (including viruses), maintains the privacy of your files and information about you, and authenticates the identity of each code provider. You can subject applications and applets to security controls when you need to.

  • Developers:

    You can use API methods to incorporate security functionality into your programs, including cryptography services and security checks. The API framework enables you to define and integrate your own permissions (controlling access to specific resources), cryptography service implementations, security manager implementations, and policy implementations. In addition, classes are provided for management of your public/private key pairs and public key certificates from people you trust.

  • Systems administrators, developers, and users:
    JDK tools manage your keystore (database of keys and certificates); generate digital signatures for JAR files, and verify the authenticity of such signatures and the integrity of the signed contents; and create and modify the policy files that define your installation's security policy.

Observe Applet Restrictions

The Java Plug-in uses a Security Manager to keep viruses from accessing your computer through an applet. No unsigned applet is allowed to access a resource unless the Security Manager finds that permission has been explicitly granted to access that system resource. That permission is granted by an entry in a policy file.

Here's the source code for an applet named WriteFile that tries to create and to write to a file named writetest in the current directory. This applet will not be able to create the file unless it has explicit permission in a policy file.

Type this command in your command window:

appletviewer http://java.sun.com/docs/books/tutorial/security/tour1/examples/WriteFile.html
Type this command on a single line, without spaces in the URL.

You should see a message about a security exception, as shown in the following figure. This is the expected behavior; the system caught the applet trying to access a resource it does not have permission to access.

WriteFile doesn't have permission to write to writetest


Set up a Policy File to Grant the Required Permission

A policy file is an ASCII text file and can be composed via a text editor or the graphical Policy Tool utility demonstrated in this section. The Policy Tool saves you typing and eliminates the need for you to know the required syntax of policy files, thus reducing errors.

This lesson uses the Policy Tool to create a policy file named mypolicy, in which you will add a policy entry that grants code from the directory where WriteFile.class is stored permission to write the writetest file.

Follow these steps to create and modify your new policy file:

  1. Start Policy Tool

  2. Grant the Required Permission

  3. Save the Policy File


Note for UNIX Users: The steps illustrate creating the policy file for a Windows system. The steps are exactly the same if you are working on a UNIX system. Where the text says to store the policy file in the C:\Test directory, you can store it in another directory. The examples in the step See the Policy File Effects and in the lesson Quick Tour of Controlling Applications assume that you stored it in the ~/test directory.

Start Policy Tool
To start Policy Tool, simply type the following at the command line:
policytool

This brings up the Policy Tool window.

Whenever Policy Tool is started, it attempts to fill in this window with policy information from the user policy file. The user policy file is named .java.policy by default in your home directory. If Policy Tool cannot find the user policy file, it issues a warning and displays a blank Policy Tool window (a window with headings and buttons but no data in it), as shown in the following figure.


This figure has been reduced to fit on the page.
Click the image to view it at its natural size.

You can then proceed to either open an existing policy file or to create a new policy file.

The first time you run the Policy Tool, you will see the blank Policy Tool window, since a user policy file does not yet exist. You can immediately proceed to create a new policy file, as described in the next step.


Grant the Required Permission
To grant the WriteFile applet permission to create and to write to the writetest file, you must create a policy entry granting this permission. To create a new entry, click on the Add Policy Entry button in the main Policy Tool window. This displays the Policy Entry dialog box as shown in the following figure.

This figure has been reduced to fit on the page.
Click the image to view it at its natural size.

A policy entry specifies one or more permissions for code from a particular code source - code from a particular location (URL), code signed by a particular entity, or both.

The CodeBase and the SignedBy text boxes specify which code you want to grant the permission(s) you will be adding in the file.

  • A CodeBase value indicates the code source location; you grant the permission(s) to code from that location. An empty CodeBase entry signifies "any code" -- it does not matter where the code originates.

  • A SignedBy value indicates the alias for a certificate stored in a keystore. The public key within that certificate is used to verify the digital signature on the code. You grant permission to code signed by the private key corresponding to the public key in the keystore entry specified by the alias. The SignedBy entry is optional; omitting it signifies "any signer" -- it does not matter whether the code is signed, or by whom.

If you have both a CodeBase and a SignedBy entry, the permission(s) are granted only to code that is both from the specified location and signed by the named alias.

To grant WriteFile the permission it needs, you can grant permission to all code from the location (URL) where WriteFile.class is stored.

Type the following URL into the CodeBase text box of the Policy Entry dialog box:

http://java.sun.com/docs/books/tutorial/security/tour1/examples/
Note: This is a URL. Therefore, it must always use slashes as separators, not backslashes.

Leave the SignedBy text box blank, since you aren't requiring the code to be signed.


Note: To grant the permission to any code (.class file) not just from the directory specified previously but from the security directory and its subdirectories, type the following URL into the CodeBase box:
http://java.sun.com/docs/books/tutorial/security/-

You have specified where the code comes from (the CodeBase), and that the code does not have to be signed (since there is no SignedBy value). Now you are ready to grant permissions to that code.

Click on the Add Permission button to display the Permissions dialog box.


This figure has been reduced to fit on the page.
Click the image to view it at its natural size.
Follow these steps to grant code from the specified CodeBase permission to write (and thus also to create) the file named writetest.
  1. Choose File Permission from the Permission drop-down list. The complete permission type name (java.io.FilePermission) now displays in the text box to the right of the drop-down list.

  2. Type the following in the text box to the right of the list labeled Target Name to specify the file named writetest:
    writetest
  3. Specify write access by choosing the write option from the Actions drop-down list.
The Permissions dialog box now looks like the following.

This figure has been reduced to fit on the page.
Click the image to view it at its natural size.

Click on the OK button. The new permission displays in a line in the Policy Entry dialog. So now the policy entry window looks like this.


This figure has been reduced to fit on the page.
Click the image to view it at its natural size.

You have now specified this policy entry, so click on the Done button in the Policy Entry dialog. The Policy Tool window now contains a line representing the policy entry, showing the CodeBase value.


This figure has been reduced to fit on the page.
Click the image to view it at its natural size.


Save the Policy File

To save the new policy file you've been creating, choose the Save As command from the File menu. This displays the Save As dialog box.

The examples in this lesson and in the Quick Tour of Controlling Applications lesson assume that you stored the policy file in the Test directory on the C: drive.

Navigate to the Test directory. Type the file name mypolicy and click on Save.

The policy file is now saved, and its name and path are shown in the text box labeled Policy File.


This figure has been reduced to fit on the page.
Click the image to view it at its natural size.

Exit Policy Tool by choosing Exit from the File menu.


See the Policy File Effects
Now that you have created the mypolicy policy file, you can execute the WriteFile applet to create and to write the file writetest, as shown in the following figure.
WriteFile can now access writetest

Whenever you run an applet, or an application with a security manager, the policy files that are loaded and used by default are the ones specified in the "security properties file", which is located in one of the following directories:

  Windows:
java.home\lib\security\java.security
UNIX:
java.home/lib/security/java.security
Note: The java.home environment variable names the directory into which the JRE was installed.

The policy file locations are specified as the values of properties whose names take the form

policy.url.n
Where n indicates a number. Specify each property value in a line that takes the following form:
policy.url.n=URL
Where URL is a URL specification. For example, the default policy files, sometimes referred to as the system and user policy files, respectively, are defined in the security properties file as
policy.url.1=file:${java.home}/lib/security/java.policy
policy.url.2=file:${user.home}/.java.policy

Note: Use of the notation ${propName} in the security properties file is a way of specifying the value of a property. Thus ${java.home} will be replaced at runtime by the actual value of the "java.home" property, which indicates the directory into which the JRE was installed, and ${user.home} will be replaced by the value of the "user.home" property, for example, C:\Windows.
In the previous step you did not modify one of these existing policy files. You created a new policy file named mypolicy. There are two possible ways you can have the mypolicy file be considered as part of the overall policy, in addition to the policy files specified in the security properties file. You can either specify the additional policy file in a property passed to the runtime system, as described in Approach 1, or add a line in the security properties file specifying the additional policy file, as described in Approach 2.

Note: On a UNIX system, you must have DNS configured in order for the WriteFile program to be downloaded from the public web site, shown in the command below. You need to have dns in the list of lookup services for hosts in your /etc/nsswitch.conf file, as in
    hosts:    dns files nis
You also need a /etc/resolv.conf file with a list of nameservers. Consult your system administrator for more information.

Approach 1

You can use the appletviewer command-line argument, -J-Djava.security.policy, to specify a policy file that should be used, in addition to the ones specified in the security properties file. To run the WriteFile applet with your new mypolicy policy file included, type the following in the directory in which mypolicy is stored:
appletviewer -J-Djava.security.policy=mypolicy 
http://java.sun.com/docs/books/tutorial/security/tour1/examples/WriteFile.html

Notes:
  • Type this command as a single line, with a space between mypolicy and the URL, and no spaces in the URL. Multiple lines are used in this example for legibility purposes.

  • If this command line is longer than the maximum number of characters you are allowed to type on a single line, do the following. Create and save a text file containing the full command, and name the file with a .bat extension, for example, wf.bat. Then in your command window, type the name of the .bat file instead of the command.

If the applet still reports an error, you must troubleshoot the policy file. Use the Policy Tool to open the mypolicy file (using File > Open) and check the policy entries you just created in the previous step, Set Up a Policy File to Grant the Required Permissions.

To view or edit an existing policy entry, click on the line displaying that entry in the main Policy Tool window, then choose the Edit Policy Entry button. You can also double-click the line for that entry.

This launches the same type of Policy Entry dialog box that displays when you are adding a new policy entry after choosing the Add Policy Entry button, except in this case the dialog box is filled in with the existing policy entry information. To change the information, retype it (for the CodeBase and SignedBy values) or add, remove, or modify permissions.

Approach 2

You can specify a number of URLs (including ones of the form "http://") in policy.url.n properties in the security properties file, and all the designated policy files will get loaded.

So one way to have our mypolicy file's policy entry considered by the appletviewer is to add an entry specifying that policy file in the security properties file.


Important: If you are running your own copy of the JDK, you can easily edit your security properties file. If you are running a version shared with other users, you may only be able to modify the system-wide security properties file if you have write access to it or if you ask your system administrator to modify the file when appropriate. However, it's probably not appropriate for you to make modifications to a system-wide policy file for this tutorial test. We suggest that you just read the following to see how it is done or that you install your own private version of the JDK to use for the tutorial lessons.

To modify the security properties file, open it in an editor suitable for editing an ASCII text file. Then add the following line after the line starting with policy.url.2:

  Windows:
policy.url.3=file:/C:/Test/mypolicy
UNIX:
policy.url.3=file:${user.home}/test/mypolicy

On a UNIX system you can also explicitly specify your home directory:

policy.url.3=file:/home/susanj/test/mypolicy

Now you can run the following:

appletviewer http://java.sun.com/docs/books/tutorial/
security1.2/tour1/examples/WriteFile.html

Type this command on one line, without spaces in the URL.

If you still get a security exception, you must troubleshoot your new policy file. Use the Policy Tool to check the policy entry you just created in the previous step, Set Up a Policy File to Grant the Required Permissions. Change any typos or other errors.


Important: The mypolicy policy file is also used in the Quick Tour of Controlling Applications lesson. You do not need to include the mypolicy file unless you are running this Tutorial lesson. To exclude this file, open the security properties file and delete the line you just added.

Thursday, May 21, 2009

Semantic Web - Microformat

A microformat is a web-based[1] approach to semantic markup that seeks to re-use existing XHTML and HTML tags to convey metadata[2] and other attributes. This approach allows information intended for end-users (such as contact information, geographic coordinates, calendar events, and the like) to also be automatically processed by software.

Although the content of web pages is technically already capable of "automated processing", and has been since the inception of the web, such processing is difficult because the traditional markup tags used to display information on the web do not describe what the information means.[3] Microformats are intended to bridge this gap by attaching semantics, and thereby obviate other, more complicated methods of automated processing, such as natural language processing or screen scraping. The use, adoption and processing of microformats enables data items to be indexed, searched for, saved or cross-referenced, so that information can be reused or combined.[3]

Current microformats allow the encoding and extraction of events, contact information, social relationships and so on. More are being developed. Version 3 of the Firefox browser,[4] as well as version 8 of Internet Explorer[5] are expected to include native support for microformats.

Background

Microformats emerged as part of a grassroots movement to make recognizable data items (such as events, contact details or geographical locations) capable of automated processing by software, as well as directly readable by end-users.[3][6] Link-based microformats emerged first. These include vote links that express opinions of the linked page, which can be tallied into instant polls by search engines.[7]

As the microformats community grew, CommerceNet, a nonprofit organization that promotes electronic commerce on the Internet, helped sponsor and promote the technology and support the microformats community in various ways.[7] CommerceNet also helped co-found the Microformats.org community site.[7]

Neither CommerceNet nor Microformats.org is a standards body. The microformats community is an open wiki, mailing list, and Internet relay chat (IRC) channel.[7] Most of the existing microformats were created at the Microformats.org wiki and associated mailing list, by a process of gathering examples of web publishing behaviour, then codifying it. Some other microformats (such as rel=nofollow and unAPI) have been proposed, or developed, elsewhere.

Technical overview

XHTML and HTML standards allow for semantics to be embedded and encoded within the attributes of markup tags. Microformats take advantage of these standards by indicating the presence of metadata using the following attributes:

  • class
  • rel
  • rev (in one case, otherwise deprecated in microformats[8])

For example, in the text "The birds roosted at 52.48,-1.89" is a pair of numbers which may be understood, from their context, to be a set of geographic coordinates. By wrapping them in spans (or other HTML elements) with specific class names (in this case geo, latitude and longitude, all part of the geo microformat specification):

The birds roosted at
<span class="geo">
<span class="latitude">52.48</span>,
<span class="longitude">-1.89</span>
</span>

Machines can be told exactly what each value represents and can then perform a variety of tasks such as indexing it, looking it up on a map and exporting it to a GPS device.

Example

In this example, the contact information is presented as follows:

 <div>
<div>Joe Doe</div>
<div>The Example Company</div>
<div>604-555-1234</div>
<a href="http://example.com/">http://example.com/</a>
</div>

With hCard microformat markup, that becomes:

 <div class="vcard">
<div class="fn">Joe Doe</div>
<div class="org">The Example Company</div>
<div class="tel">604-555-1234</div>
<a class="url" href="http://example.com/">http://example.com/</a>
</div>

Here, the formatted name (fn), organisation (org), telephone number (tel) and web address (url) have been identified using specific class names and the whole thing is wrapped in class="vcard", which indicates that the other classes form an hCard (short for "HTML vCard") and are not merely coincidentally named. Other, optional, hCard classes also exist. It is now possible for software, such as browser plug-ins, to extract the information, and transfer it to other applications, such as an address book.

In-context examples

For annotated examples of microformats on live pages, see HCard#Live example and Geo (microformat)#Three_classes.

Specific microformats

Several microformats have been developed to enable semantic markup of particular types of information.

  • hAtom - for marking up Atom feeds from within standard HTML
  • hCalendar - for events
  • hCard - for contact information; includes:

Microformats under development

Among the many proposed microformats[13], the following are undergoing active development:

  • hAudio - for audio files and references to released recordings
  • hRecipe [14]
  • citation - for citing references
  • currency - for amounts of money
  • figure - for associating captions with images [15]
  • geo extensions - for places on Mars, the Moon, and other such bodies; for altitude; and for collections of waypoints marking routes or boundaries
  • species - For the names of living things.
  • measure - For physical quantities, structured data-values.[16]

Uses of microformats

Using microformats within HTML code provides additional formatting and semantic data that can be used by applications. These could be applications that collect data about on-line resources, such as web crawlers, or desktop applications such as e-mail clients or scheduling software. They can also be used to facilitate "mash ups" such as exporting all of the geographical locations on a web page into Google Maps, to visualize them spatially.

Several browser extensions, such as Operator for Firefox and Oomph for Internet Explorer, provide the ability to detect microformats within an HTML document and export them into formats compatible with contact management and calendar utilities, such as Microsoft Outlook. Yahoo! Query Language can be used to extract microformats from web pages.[17]

Microsoft expressed a desire to incorporate Microformats into upcoming projects;[18] as have other software companies.

In Wikipedia - and more generally in MediaWiki - microformats are used as part of templates like {{coord}}.

Alex Faaborg summarizes the arguments for putting the responsibility for microformat user interfaces in the web browser rather than making more complicated HTML:[19]

  • Only the web browser knows what applications are accessible to the user and what the user's preferences are
  • It lowers the barrier to entry for web site developers if they only need to do the markup and not handle "appearance" or "action" issues
  • Retains backwards compatibility with web browsers that don't support microformats
  • The web browser presents a single point of entry from the web to the user's computer, which simplifies security issues

Semantic Web

The Semantic Web is an evolving extension of the World Wide Web in which the semantics of information and services on the web is defined, making it possible for the web to understand and satisfy the requests of people and machines to use the web content.[1][2] It derives from World Wide Web Consortium director Sir Tim Berners-Lee's vision of the Web as a universal medium for data, information, and knowledge exchange.[3]

At its core, the semantic web comprises a set of design principles,[4] collaborative working groups, and a variety of enabling technologies. Some elements of the semantic web are expressed as prospective future possibilities that are yet to be implemented or realized.[2] Other elements of the semantic web are expressed in formal specifications.[5] Some of these include Resource Description Framework (RDF), a variety of data interchange formats (e.g. RDF/XML, N3, Turtle, N-Triples), and notations such as RDF Schema (RDFS) and the Web Ontology Language (OWL), all of which are intended to provide a formal description of concepts, terms, and relationships within a given knowledge domain.


Purpose

Humans are capable of using the Web to carry out tasks such as finding the Finnish word for "monkey", reserving a library book, and searching for a low price for a DVD. However, a computer cannot accomplish the same tasks without human direction because web pages are designed to be read by people, not machines. The semantic web is a vision of information that is understandable by computers, so that they can perform more of the tedious work involved in finding, sharing, and combining information on the web.

Limitations of HTML


Many files on a typical computer can be loosely divided into documents and data. Documents like mail messages, reports, and brochures are read by humans. Data, like calendars, addressbooks, playlists, and spreadsheets are presented using an application program which lets them be viewed, searched and combined in many ways.

Currently, the World Wide Web is based mainly on documents written in Hypertext Markup Language (HTML), a markup convention that is used for coding a body of text interspersed with multimedia objects such as images and interactive forms. Metadata tags, for example

<meta name="keywords" content="computing, computer studies, computer">
<meta name="description" content="Cheap widgets for sale">
<meta name="author" content="John Doe">

provide a method by which computers can categorise the content of web pages.

With HTML and a tool to render it (perhaps web browser software, perhaps another user agent), one can create and present a page that lists items for sale. The HTML of this catalog page can make simple, document-level assertions such as "this document's title is 'Widget Superstore'", but there is no capability within the HTML itself to assert unambiguously that, for example, item number X586172 is an Acme Gizmo with a retail price of €199, or that it is a consumer product. Rather, HTML can only say that the span of text "X586172" is something that should be positioned near "Acme Gizmo" and "€ 199", etc. There is no way to say "this is a catalog" or even to establish that "Acme Gizmo" is a kind of title or that "€ 199" is a price. There is also no way to express that these pieces of information are bound together in describing a discrete item, distinct from other items perhaps listed on the page.

Semantic HTML refers to the traditional HTML practice of markup following intention, rather than specifying layout details directly. For example, the use of <em> denoting "emphasis" rather than <i>, which specifies italics. Layout details are left up to the browser, in combination with Cascading Style Sheets. But this practice falls short of specifying the semantics of objects such as items for sale or prices.

Microformats represent unofficial attempts to extend HTML syntax to create machine-readable semantic markup about objects such as retail stores and items for sale.

Semantic Web solutions


The Semantic Web takes the solution further. It involves publishing in languages specifically designed for data: Resource Description Framework (RDF), Web Ontology Language (OWL), and Extensible Markup Language (XML). HTML describes documents and the links between them. RDF, OWL, and XML, by contrast, can describe arbitrary things such as people, meetings, or airplane parts. Tim Berners-Lee calls the resulting network of Linked Data the Giant Global Graph, in contrast to the HTML-based World Wide Web.

These technologies are combined in order to provide descriptions that supplement or replace the content of Web documents. Thus, content may manifest as descriptive data stored in Web-accessible databases, or as markup within documents (particularly, in Extensible HTML (XHTML) interspersed with XML, or, more often, purely in XML, with layout or rendering cues stored separately). The machine-readable descriptions enable content managers to add meaning to the content, i.e. to describe the structure of the knowledge we have about that content. In this way, a machine can process knowledge itself, instead of text, using processes similar to human deductive reasoning and inference, thereby obtaining more meaningful results and helping computers to perform automated information gathering and research.

An example of a tag that would be used in a non-semantic web page:

<item>cat</item>

Encoding similar information in a semantic web page might look like this:

<item rdf:about="http://dbpedia.org/resource/Cat">Cat</item>

Saturday, May 16, 2009

Creating and Parsing JSON data

JSON (JavaScript Object Notation) is a lightweight computer data interchange format. It is a text-based, human-readable format for representing simple data structures and associative arrays (called objects). The JSON format is specified in RFC 4627 by Douglas Crockford. The official Internet media type for JSON is application/json.

The JSON format is often used for transmitting structured data over a network connection in a process called serialization. Its main application is in AJAX web application programming, where it serves as an alternative to the traditional use of the XML format.

Supported data types

  1. Number (integer, real, or floating point)
  2. String (double-quoted Unicode with backslash escapement)
  3. Boolean (true and false)
  4. Array (an ordered sequence of values, comma-separated and enclosed in square brackets)
  5. Object (collection of key/value pairs, comma-separated and enclosed in curly brackets)
  6. null

Syntax

The following example shows the JSON representation of an object that describes a person. The object has string fields for first name and last name, contains an object representing the person’s address, and contains a list of phone numbers (an array).

  1. {
  2. "firstName": "John",
  3. "lastName": "Smith",
  4. "address": {
  5. "streetAddress": "21 2nd Street",
  6. "city": "New York",
  7. "state": "NY",
  8. "postalCode": 10021
  9. },
  10. "phoneNumbers": [
  11. "212 732-1234",
  12. "646 123-4567"
  13. ]
  14. }

Creating JSON data in Java

JSON.org has provided libraries to create/parse JSON data through Java code. These libraries can be used in any Java/J2EE project including Servlet, Struts, JSF, JSP etc and JSON data can be created.

Download JAR file json-rpc-1.0.jar (75 kb)

Use JSONObject class to create JSON data in Java. A JSONObject is an unordered collection of name/value pairs. Its external form is a string wrapped in curly braces with colons between the names and values, and commas between the values and names. The internal form is an object having get() and opt() methods for accessing the values by name, and put() methods for adding or replacing values by name. The values can be any of these types: Boolean, JSONArray, JSONObject, Number, and String, or the JSONObject.NULL object.

  1. import org.json.JSONObject;
  2. ...
  3. ...
  4. JSONObject json = new JSONObject();
  5. json.put("city", "Mumbai");
  6. json.put("country", "India");
  7. ...
  8. String output = json.toString();
  9. ...

Thus by using toString() method you can get the output in JSON format.

JSON Array in Java

A JSONArray is an ordered sequence of values. Its external text form is a string wrapped in square brackets with commas separating the values. The internal form is an object having get and opt methods for accessing the values by index, and put methods for adding or replacing values. The values can be any of these types: Boolean, JSONArray, JSONObject, Number, String, or the JSONObject.NULL object.

The constructor can convert a JSON text into a Java object. The toString method converts to JSON text.

JSONArray class can also be used to convert a collection of Java beans into JSON data. Similar to JSONObject, JSONArray has a put() method that can be used to put a collection into JSON object.

Thus by using JSONArray you can handle any type of data and convert corresponding JSON output.

Useful Java Code Snippet

Convert String to Date in Java

  1. java.util.Date = java.text.DateFormat.getDateInstance().parse(date String);

or

  1. SimpleDateFormat format = new SimpleDateFormat( "dd.MM.yyyy" );
  2. Date date = format.parse( myString )

Connecting to Oracle using Java JDBC

  1. public class OracleJdbcTest
  2. {
  3. String driverClass = "oracle.jdbc.driver.OracleDriver";

  4. Connection con;

  5. public void init(FileInputStream fs) throws ClassNotFoundException, SQLException, FileNotFoundException, IOException
  6. {
  7. Properties props = new Properties();
  8. props.load(fs);
  9. String url = props.getProperty("db.url");
  10. String userName = props.getProperty("db.user");
  11. String password = props.getProperty("db.password");
  12. Class.forName(driverClass);

  13. con=DriverManager.getConnection(url, userName, password);
  14. }

  15. public void fetch() throws SQLException, IOException
  16. {
  17. PreparedStatement ps = con.prepareStatement("select SYSDATE from dual");
  18. ResultSet rs = ps.executeQuery();

  19. while (rs.next())
  20. {
  21. // do the thing you do
  22. }
  23. rs.close();
  24. ps.close();
  25. }

  26. public static void main(String[] args)
  27. {
  28. OracleJdbcTest test = new OracleJdbcTest();
  29. test.init();
  30. test.fetch();
  31. }
  32. }

Java Fast File Copy using NIO

  1. public static void fileCopy( File in, File out )
  2. throws IOException
  3. {
  4. FileChannel inChannel = new FileInputStream( in ).getChannel();
  5. FileChannel outChannel = new FileOutputStream( out ).getChannel();
  6. try
  7. {
  8. // inChannel.transferTo(0, inChannel.size(), outChannel); // original -- apparently has trouble copying large files on Windows

  9. // magic number for Windows, 64Mb - 32Kb)
  10. int maxCount = (64 * 1024 * 1024) - (32 * 1024);
  11. long size = inChannel.size();
  12. long position = 0;
  13. while ( position <>
  14. {
  15. position += inChannel.transferTo( position, maxCount, outChannel );
  16. }
  17. }
  18. finally
  19. {
  20. if ( inChannel != null )
  21. {
  22. inChannel.close();
  23. }
  24. if ( outChannel != null )
  25. {
  26. outChannel.close();
  27. }
  28. }
  29. }

PDF Generation in Java using iText JAR

Read this article for more details.

  1. import java.io.File;
  2. import java.io.FileOutputStream;
  3. import java.io.OutputStream;
  4. import java.util.Date;

  5. import com.lowagie.text.Document;
  6. import com.lowagie.text.Paragraph;
  7. import com.lowagie.text.pdf.PdfWriter;

  8. public class GeneratePDF {

  9. public static void main(String[] args) {
  10. try {
  11. OutputStream file = new FileOutputStream(new File("C:\\Test.pdf"));

  12. Document document = new Document();
  13. PdfWriter.getInstance(document, file);
  14. document.open();
  15. document.add(new Paragraph("Hello Kiran"));
  16. document.add(new Paragraph(new Date().toString()));

  17. document.close();
  18. file.close();

  19. } catch (Exception e) {

  20. e.printStackTrace();
  21. }
  22. }
  23. }

HTTP Proxy setting in Java

Read this article for more details.

  1. System.getProperties().put("http.proxyHost", "someProxyURL");
  2. System.getProperties().put("http.proxyPort", "someProxyPort");
  3. System.getProperties().put("http.proxyUser", "someUserName");
  4. System.getProperties().put("http.proxyPassword", "somePassword");

Java Singleton example

Read this article for more details.
Update: Thanks Markus for the comment. I have updated the code and changed it to more robust implementation.

  1. public class SimpleSingleton {
  2. private static SimpleSingleton singleInstance = new SimpleSingleton();

  3. //Marking default constructor private
  4. //to avoid direct instantiation.
  5. private SimpleSingleton() {
  6. }

  7. //Get instance for class SimpleSingleton
  8. public static SimpleSingleton getInstance() {

  9. return singleInstance;
  10. }
  11. }

One more implementation of Singleton class. Thanks to Ralph and Lukasz Zielinski for pointing this out.

  1. public enum SimpleSingleton {
  2. INSTANCE;
  3. public void doSomething() {
  4. }
  5. }

  6. //Call the method from Singleton:
  7. SimpleSingleton.INSTANCE.doSomething();

Capture screen shots in Java

Read this article for more details.

  1. import java.awt.Dimension;
  2. import java.awt.Rectangle;
  3. import java.awt.Robot;
  4. import java.awt.Toolkit;
  5. import java.awt.image.BufferedImage;
  6. import javax.imageio.ImageIO;
  7. import java.io.File;

  8. ...

  9. public void captureScreen(String fileName) throws Exception {

  10. Dimension screenSize = Toolkit.getDefaultToolkit().getScreenSize();
  11. Rectangle screenRectangle = new Rectangle(screenSize);
  12. Robot robot = new Robot();
  13. BufferedImage image = robot.createScreenCapture(screenRectangle);
  14. ImageIO.write(image, "png", new File(fileName));

  15. }
  16. ...

14. Files-Directory listing in Java

  1. File dir = new File("directoryName");
  2. String[] children = dir.list();
  3. if (children == null) {
  4. // Either dir does not exist or is not a directory
  5. } else {
  6. for (int i=0; i <>
  7. // Get filename of file or directory
  8. String filename = children[i];
  9. }
  10. }

  11. // It is also possible to filter the list of returned files.
  12. // This example does not return any files that start with `.'.
  13. FilenameFilter filter = new FilenameFilter() {
  14. public boolean accept(File dir, String name) {
  15. return !name.startsWith(".");
  16. }
  17. };
  18. children = dir.list(filter);

  19. // The list of files can also be retrieved as File objects
  20. File[] files = dir.listFiles();

  21. // This filter only returns directories
  22. FileFilter fileFilter = new FileFilter() {
  23. public boolean accept(File file) {
  24. return file.isDirectory();
  25. }
  26. };
  27. files = dir.listFiles(fileFilter);


Creating ZIP and JAR Files in Java

  1. import java.util.zip.*;
  2. import java.io.*;

  3. public class ZipIt {
  4. public static void main(String args[]) throws IOException {
  5. if (args.length < 2) {
  6. System.err.println("usage: java ZipIt Zip.zip file1 file2 file3");
  7. System.exit(-1);
  8. }
  9. File zipFile = new File(args[0]);
  10. if (zipFile.exists()) {
  11. System.err.println("Zip file already exists, please try another");
  12. System.exit(-2);
  13. }
  14. FileOutputStream fos = new FileOutputStream(zipFile);
  15. ZipOutputStream zos = new ZipOutputStream(fos);
  16. int bytesRead;
  17. byte[] buffer = new byte[1024];
  18. CRC32 crc = new CRC32();
  19. for (int i=1, n=args.length; i <>
  20. String name = args[i];
  21. File file = new File(name);
  22. if (!file.exists()) {
  23. System.err.println("Skipping: " + name);
  24. continue;
  25. }
  26. BufferedInputStream bis = new BufferedInputStream(
  27. new FileInputStream(file));
  28. crc.reset();
  29. while ((bytesRead = bis.read(buffer)) != -1) {
  30. crc.update(buffer, 0, bytesRead);
  31. }
  32. bis.close();
  33. // Reset to beginning of input stream
  34. bis = new BufferedInputStream(
  35. new FileInputStream(file));
  36. ZipEntry entry = new ZipEntry(name);
  37. entry.setMethod(ZipEntry.STORED);
  38. entry.setCompressedSize(file.length());
  39. entry.setSize(file.length());
  40. entry.setCrc(crc.getValue());
  41. zos.putNextEntry(entry);
  42. while ((bytesRead = bis.read(buffer)) != -1) {
  43. zos.write(buffer, 0, bytesRead);
  44. }
  45. bis.close();
  46. }
  47. zos.close();
  48. }
  49. }

Parsing / Reading XML file in Java

  1. package net.viralpatel.java.xmlparser;

  2. import java.io.File;
  3. import javax.xml.parsers.DocumentBuilder;
  4. import javax.xml.parsers.DocumentBuilderFactory;

  5. import org.w3c.dom.Document;
  6. import org.w3c.dom.Element;
  7. import org.w3c.dom.Node;
  8. import org.w3c.dom.NodeList;

  9. public class XMLParser {

  10. public void getAllUserNames(String fileName) {
  11. try {
  12. DocumentBuilderFactory dbf = DocumentBuilderFactory.newInstance();
  13. DocumentBuilder db = dbf.newDocumentBuilder();
  14. File file = new File(fileName);
  15. if (file.exists()) {
  16. Document doc = db.parse(file);
  17. Element docEle = doc.getDocumentElement();

  18. // Print root element of the document
  19. System.out.println("Root element of the document: "
  20. + docEle.getNodeName());

  21. NodeList studentList = docEle.getElementsByTagName("student");

  22. // Print total student elements in document
  23. System.out
  24. .println("Total students: " + studentList.getLength());

  25. if (studentList != null && studentList.getLength() > 0) {
  26. for (int i = 0; i <>

  27. Node node = studentList.item(i);

  28. if (node.getNodeType() == Node.ELEMENT_NODE) {

  29. System.out
  30. .println("=====================");

  31. Element e = (Element) node;
  32. NodeList nodeList = e.getElementsByTagName("name");
  33. System.out.println("Name: "
  34. + nodeList.item(0).getChildNodes().item(0)
  35. .getNodeValue());

  36. nodeList = e.getElementsByTagName("grade");
  37. System.out.println("Grade: "
  38. + nodeList.item(0).getChildNodes().item(0)
  39. .getNodeValue());

  40. nodeList = e.getElementsByTagName("age");
  41. System.out.println("Age: "
  42. + nodeList.item(0).getChildNodes().item(0)
  43. .getNodeValue());
  44. }
  45. }
  46. } else {
  47. System.exit(1);
  48. }
  49. }
  50. } catch (Exception e) {
  51. System.out.println(e);
  52. }
  53. }
  54. public static void main(String[] args) {

  55. XMLParser parser = new XMLParser();
  56. parser.getAllUserNames("c:\\test.xml");
  57. }
  58. }

Convert Array to Map in Java

  1. import java.util.Map;
  2. import org.apache.commons.lang.ArrayUtils;

  3. public class Main {

  4. public static void main(String[] args) {
  5. String[][] countries = { { "United States", "New York" }, { "United Kingdom", "London" },
  6. { "Netherland", "Amsterdam" }, { "Japan", "Tokyo" }, { "France", "Paris" } };

  7. Map countryCapitals = ArrayUtils.toMap(countries);

  8. System.out.println("Capital of Japan is " + countryCapitals.get("Japan"));
  9. System.out.println("Capital of France is " + countryCapitals.get("France"));
  10. }
  11. }

Send Email using Java

  1. import javax.mail.*;
  2. import javax.mail.internet.*;
  3. import java.util.*;

  4. public void postMail( String recipients[ ], String subject, String message , String from) throws MessagingException
  5. {
  6. boolean debug = false;

  7. //Set the host smtp address
  8. Properties props = new Properties();
  9. props.put("mail.smtp.host", "smtp.example.com");

  10. // create some properties and get the default Session
  11. Session session = Session.getDefaultInstance(props, null);
  12. session.setDebug(debug);

  13. // create a message
  14. Message msg = new MimeMessage(session);

  15. // set the from and to address
  16. InternetAddress addressFrom = new InternetAddress(from);
  17. msg.setFrom(addressFrom);

  18. InternetAddress[] addressTo = new InternetAddress[recipients.length];
  19. for (int i = 0; i <>
  20. {
  21. addressTo[i] = new InternetAddress(recipients[i]);
  22. }
  23. msg.setRecipients(Message.RecipientType.TO, addressTo);

  24. // Optional : You can also set your custom headers in the Email if you Want
  25. msg.addHeader("MyHeaderName", "myHeaderValue");

  26. // Setting the Subject and Content Type
  27. msg.setSubject(subject);
  28. msg.setContent(message, "text/plain");
  29. Transport.send(msg);
  30. }



Send HTTP request & fetching data using Java

  1. import java.io.BufferedReader;
  2. import java.io.InputStreamReader;
  3. import java.net.URL;
  4. public class Main {
  5. public static void main(String[] args) {
  6. try {
  7. URL my_url = new URL("http://www.viralpatel.net/blogs/");
  8. BufferedReader br = new BufferedReader(new InputStreamReader(my_url.openStream()));
  9. String strTemp = "";
  10. while(null != (strTemp = br.readLine())){
  11. System.out.println(strTemp);
  12. }
  13. } catch (Exception ex) {
  14. ex.printStackTrace();
  15. }
  16. }
  17. }


Resize an Array in Java

  1. /**
  2. * Reallocates an array with a new size, and copies the contents
  3. * of the old array to the new array.
  4. * @param oldArray the old array, to be reallocated.
  5. * @param newSize the new array size.
  6. * @return A new array with the same contents.
  7. */
  8. private static Object resizeArray (Object oldArray, int newSize) {
  9. int oldSize = java.lang.reflect.Array.getLength(oldArray);
  10. Class elementType = oldArray.getClass().getComponentType();
  11. Object newArray = java.lang.reflect.Array.newInstance(
  12. elementType,newSize);
  13. int preserveLength = Math.min(oldSize,newSize);
  14. if (preserveLength > 0)
  15. System.arraycopy (oldArray,0,newArray,0,preserveLength);
  16. return newArray;
  17. }
  18. // Test routine for resizeArray().
  19. public static void main (String[] args) {
  20. int[] a = {1,2,3};
  21. a = (int[])resizeArray(a,5);
  22. a[3] = 4;
  23. a[4] = 5;
  24. for (int i=0; i
  25. System.out.println (a[i]);
  26. }