Friday, May 22, 2009

Experience with using multiple datasources in JEE application

We started with multiple data sources in one of our project. One of the reasons why it was decided to use multiple data sources was to reduce performance impact. It was SOA so we have started assigning a dedicated data source with dedicated authorization id to each service. The very first problem what we encountered here was maintenance of authorization ids. As different data sources were using different authorization id, we had to keep track of what data source uses which authorization id. After having around 8-10 data sources this maintenance became bizarre. We learn from that experience and made all data sources using same authorization id in a given environment. So now we have all data sources using same id in DEV environment and other id into QA environment and all. This model works pretty well with us.

Second problem in using multiple data sources was mysterious dead lock issue . We were using all these data sources for accessing same DB. When multiple data sources are nvolved into single logical unit of work (i.e. logical transaction), each data source starts its own transactional branch. If this transaction with isolation level set to READ_COMMITTED involves multiple tables and some of them are having foreign key relationship then it might happen that one transaction branch insert a record into table x and other transaction branch also tries to insert a record into table y which has foreign key relationship to table x. In this scenario table y will time out if record to be inserted into table y is related to a record inserted into table x. Reason behind this is, record inserted into table x is locked by transaction A so record to be inserted into table y can not read the record from table x for checking referential integrity as its in other transaction and isolation level is set to READ_COMMITTED.

So important points to be considered in using multiple data sources are:
  1. Please try to stick to same authorization id for a given environment, if not required.
  2. Multiple data sources should not be pointing to same database, otherwise it might lead to dead lock issue quite often. Use multiple data sources if its required to point to different databases.
  3. In case of multiple data source and remote EJB call scenario one should use XA data source as it might lead to a XA transaction.

Sunday, February 15, 2009

My study on using stored procedures in JEE applications

Pros:
  1. Good at performance as all operations are happens inside DBMS system itself. Also nowadays most of DBMS system has mechanism of caching query execution plan so it can perform even faster at future calls. Query execution plan remains same during each round of executions of a stored procedure.
  2. Provides another layer of abstraction by providing separation between application and underlying DBMS system.
  3. Reduces network traffic by avoiding repetitive calls to DBMS system.
Cons:
  1. There are very high chances that business logic gets spread at two places. So it’s difficult to manage the business logic.
  2. We need versioning of stored procedures as its part of project deliverable and we can not check out a copy of stored procedure from DBMS system. We are left with two copies, one which is present into DBMS system and the other one which is present at SVN. So it’s hard to maintain latest valid copy of a stored procedure at DBMS system.
  3. Stored procedures are bad at scalability. We can do horizontal as well as vertical scaling of J2EE application server for doing load balancing but we can not do the same for DBMS system as they don’t support clustering. Oracle has some support for clustering. So stored procedure execution can become bottleneck into multi user system with high load.
  4. Using stored procedure means another layer of abstraction and another language to learn.
  5. Hard to debug from the same IDE which is used for development. So its affects productivity.
  6. There are very high chance of java developers writing improper stored procedure like stored procedure with less efficient queries, open cursors, improper error handling, improper state management etc.
  7. Stored procedure language is procedural language. So it’s very difficult to write a complex logic into it. Even if we write it ships with all disadvantages of procedural language. E.g. If we need to do two operations on a Customer then we need to write down two procedures. Lots of code used in these two stored procedures may be same but we can not reuse it. So we might end up in having lots of code duplication.
  8. Reduces testability of application.

My comments:
Each and every person in team should understand clear distinction between data modeling operation and business operation for a given functionality. So business logic should go into service layer and the data related operation if it’s too complex then should go into stored procedures. I believe that if we have very big DBMS schema with large number of records in it, we should keep data related operation (E.g. data formatting, data aggregation etc.) into stored procedures. We should have proper process in place for managing this kind clear distinction from the starting of the project itself otherwise it might become very difficult to control it at later point of time.

Most of the above cons can be overcome by having proper processes in place as well as better tooling support.

Personally I don’t like to use stored procedures, if I have other good options at Java space for given requirements. But if stored procedure is the only good option which can fit well to the given requirements then I don’t mind in doing that keeping above distinctions in mind.

I may be wrong in stating above comment. I would like to correct my perceptions if they are wrong. It will really help me going further. Thanks in advance for the same.

Some quotes:
"If you want portability keep logic out of SQL." - Martin Fowler

"Do not use stored procedures to implement business logic. This should be done in Java business objects." - Rod Johnson

References:
http://codebetter.com/blogs/jeremy.miller/archive/2005/07/06/130094.aspx
http://www.martinfowler.com/articles/dblogic.html
http://thinkoracle.blogspot.com/2006/01/plsql-vs-j2ee.html
http://www.codinghorror.com/blog/archives/000117.html
http://www.theserverside.com/discussions/thread.tss?thread_id=32141

Tuesday, October 21, 2008

What is optimization?

As per wikipedia, "Optimization (mathematics), trying to find maxima and minima of a function".

So as per my understanding optimization is basically a search problem/function. For me optimization is the process of searching the maxima or minima into the given n-dimensional space where search is driven by the constraints imposed by the problem definition and/or environment.

Thursday, September 4, 2008

ACO Vs. PSO. What to use when?

Lot of literature is available introducing ACO and PSO. At very first time when I read them, I thought that they are different ways of solving same problems. I mean I was thinking that, they can be applied on same set of problems. Then I ran into another thought, "Then why do we need both of them?". For getting answer of this question, I read their introductions again. I got my answer. The answer is:

ACO is mostly used for discrete optimization problems.
PSO is mostly used for continuous optimization problems.

Reason behind this is:
ACO is driven by two parameters : heuristic value and pheromone value. Mostly these values are derived from parameters having discrete values.
PSO is driven by neighbor's velocity. Velocity is continuous parameter. As one of the parameters used for deriving velocity is time and time is continuous.

ACO fits better at graph searching problems while PSO fits better at ANN learning as well as pattern recognition. Reason behind this is, parameters used for graph searching are mostly discrete parameters while parameters used for learning/recognition are continuous parameters.

Friday, July 18, 2008

Can we apply Ant Colony Optmization on Tic-Tac-Toe?

Question:
I need to implement ACO in tic-tac-toc game playing. The problem is when i studied the algorithm that was on static graph called TSP problem. But in T-T-T there will be dynamic graph or tree.

My Answer:
You can always apply Ant Colony Optimization(ACO) on any search graph. Actually basic fundamental behind this concept is the Ant's foraging nature. Path from nest to food is never a static search graph. When Ant starts searching for food from its nest, it even don't know where the food is?

I have applied ACO on routing in Mobile Adhoc Network(MANETs). MANETs by nature are dynamic so its search graph is always dynamic.

For solving this problem, you need to know best heuristic for solving Tic-Tac-Toe. Once that is known to you, I can really help you out in solving your problem.

One heuristic for Tic-Tac-Toe game is the position. If you are at the centre position then you have highest chances to win(Say 3 points). Then if you are at corner position then you have lesser chances to win(Say 2 points). If you are at any other postion then you have even lesser chances to win(Say 1 point). So a heuristic you can apply is, try to go for the postion which does not make you loose as well as gives highest points. As you have at max 4 corner postions so in that case your selection of any of the available corner postion should be driven by Pheromone. I mean the postion which has highest pheromone value should be considered. So this is the postion which is good as per heurustic values and better as per prior experience(having higher pheromone intensity). I guess Tic-Tac-Toe is solved using ACO. Isn't it???

You can always apply other better Heurisitics. Also you can try to apply different weightage on Heuristics and Pheromone by changing values of Alpha and Beta.

Hope you will be in a postion to start your implementation. You are very welcome to ask your doubts here, if any.

Saturday, June 14, 2008

Ant Colony Optimization(ACO)

Introduction of Ant Colony Optimization(ACO):
Ant Colony Optimization is a probabilistic technique for searching based on Ant's foraging behavior. The concept was derived by Macro Dorigo. It best used for solving combinatorial optimization problems. Best suitable for minimization of multiple paths from a search graph.

Example applications are: TSP, Intelligent Routing, Classification etc.

Usage of Comparable interface in core Java:

Comparable interface is used to make your object comparable to other objects. Comparable interface has int compareTo(Object object). One can override this method to provide appropriate comparison of this object to the other object which is sent as an argument to this method.

Example:
Class Empoyee impements Comparable{

String name;

public int compareTo(Object otherObject){
int retValue = -1;
if(otherObject instanceOf Employee){
retValue = getName().compareTo(otherObject .getName());
}

return retValue;
}

}