Friday, July 20, 2007

Analysis of the crash...

I've been looking at the video now in step-mode. It's poor quality, but I can derive a couple of things that question the statement that the reverse thrusters had been in operation the whole time. Looking at the video and taking some screenshots however, I'm not so convinced this was the case.

(click on the pictures to see them larger)

Here are some considerations of mine that suggest a different sequence of events that seems to be a closer account:

http://www.youtube.com/watch?v=k6lO-eig_i0&mode=related&search=

Here we see the approach of the aircraft, right, where the plane sees the camera at an angle of about 300 degrees, head-up:


This image shows how the aircraft is passing the camera. The rain should be an indication of full-thrust working. The speed of the air at thrust surpasses the speed of the airplane by far. I see no waterspray being pushed in front of the airplane. Also, the size of the spray and its shape seem to indicate that the jet being pushed backwards is solely attributable to the wheels. There isn't even, as far as I can see, a buildup around the wings that indicate any kind of reverse-thrust to stop the plane at this point. Note that at this point the plane is probably around half the length of the runway. The Airbus A320 requires at least 300-500 meters of runway beyond the 1.9km that this runway has:


Here we see another camera recording the event. The plane is in view lower right and is just entering the camera frustrum:


Camera 10 has a detail view. Now we do see a waterspray being pushed forward even beyond the forward wheel. The buildup of spray around the body of the aircraft as a whole is noticeable. This is what you'd probably have to see in the first image (albeit the shape of the water spray would be longer due to higher speed), but at least the water should move higher and around the body of the aircraft if reverse-thrust is to be engaged. See that white cloud in this picture? I didn't see that in the previous pictures. Common sense tells me that if reverse thrust was working earlier, we should at least see a visible deceleration and similar upward-moving waterspray in the previous pictures.


Here we see another image a second later. In the overall movement of things, I don't see the aircraft noticeably slowing down, but it's as if it suddenly starts rolling out ( no thrust applied whatsoever ) and it's as if the spray of the reverse thrust is suddenly diminishing significantly. Has the pilot just decided to abort the landing at this point in an attempt to take off again with the remaining speed? This is probably about 3/4 down the runway or so:


This image here shows another camera about 3 seconds later. There seems to be a short flash at the left turbine, which may indicate the thruster reversing again into forward mode. I am not sure whether reversing the thrusters at this point, when still rotating reverse, would ignite a flash of some kind (maybe someone can comment?). The length of the runway in front must definitely be very limited.


Even though the news indicates that the right reversor was defective, I'm not sure whether this is truly the cause. It's very well possible that it was not operating and the aircraft logic prohibiting the operation of the left reverse thruster. However, there are other accounts of disabled reverse thrust or braking due to failure of the sensors or conditions prohibiting proper sensing of ground conditions (the A320 does not allow pilots to engage reverse-thrust in "flight" condition).

http://www.aaib.dft.gov.uk/publications/bulletins/february_2005/airbus_a320_200__c_ftdf.cfm

http://www.rvs.uni-bielefeld.de/publications/Incidents/DOCS/ComAndRep/Warsaw/warsaw-report.html

http://nakedshorts.typepad.com/nakedshorts/2005/08/debugging_airbu.html

Plus... we see that the reverse thrust does kick in at some point (picture 4), albeit much too late. If there was a full defect, this wouldn't have happened.

There seems to be quite some confusion with engineers as to how the actual braking operation works, as the manual is not very clear at this point:

http://www.pprune.org/forums/archive/index.php/t-92017.html

I can imagine that when one engine does not allow reverse thrust, the other should not apply it as this would spin the airplane around. There should be a safety mechanism in an aircraft that guarantees (more or less?) equal thrust being applied to both engines. Steering in an aircraft is not done through thrusters, but through wing action and ailerons.

Personally, if this is what happened, I can understand the stress of the pilot. There is only just enough length to land in dry weather conditions. The runway is wet. The pilot with 20 years of experience must have landed here before and know about the length of the runway. The pilot must have known about the pending maintenance action for the right turbine. The grooving has not yet been done (pending for September) due to "commercial pressure" to open the airport ( the losses would be too great ). The aircraft on touchdown does not respond to any braking commands (see potential similarity with other reports on other A320 crashes). More than halfway down the runway, the aircraft finally responds, but the length in front of the plane simply doesn't cut it. The pilots probably both decide to pull back up. The reverse-thrust is aborted and the thruster is set to forward thrust again. The speed of the aircraft at this point is far from favourable with only very little runway left. Is it possible that another 300-500m would have saved their lives? The A320 takes off and lands at around 160knots. This is 82m per second. The landing speed was above 160 knots at the time of touchdown.

Besides this plane having probably suffered a technical problem, we cannot ignore the other factors that have contributed to this. The pilots get informed that they better circle around for another landing attempt if they do not manage to land at the first 300m on the runway. Imagine the consequences if the braking system doesn't work and you're halfway down the runway (that is 8-10 seconds down). The decision you are forced to take in the next 10 seconds is crucial and every second is very, very crucial. Braking the airplane at more or less full-speed having only half of the runway left, a runway without grooves in wet conditions? This probably went through the pilot's minds at the last point of decision.

A couple of people in Brazil just put a value on the price of human life. For a jet full of people to crash in a busy airport, the monetary equivalent is the money that was made from February 2007 until now.

My expectation for any airport wherever in the world is that all international and safety norms are met, if not exceeded. THEN we can talk about weighing off extra security and safety measures versus economic benefit.

Thursday, July 19, 2007

Accountability (vs. "relaxa e morra")

I found a new, very insightful and interesting blog from Lucia Hippolito. Cientista política, historiadora e jornalista, especialista em eleições, partidos políticos e Estado brasileiro.

http://www.luciahippolito.globolog.com.br/

She also commented on the fact of lack of accountability across Brasil. With the following main observations through the text:
  • Accountability contains the idea that authority is a public servant. Elected or not, it has to be accountability for its actions to society.
  • Less stage and more debate, less uprisings and more interviews, less "law by ministry" (do other countries have this even?) and more attendance to the Congress.
Well... Accountability. Such a great word, and there is no portuguese translation! (ironic?)

When we look at the disaster of the airplane in São Paulo, some important "political processes" immediately kicked in. Nope. Not what you expect. Immediate investigations were ordered to try to blame it on the runway, but overall, the political world kept rather quite. The president has, unfortunately, not appeared on television in the last 72 hrs to send his condolences and show his commitment and compassion. Bit disappointing.

In February 2007, the airport was closed for 737, Fokker 100 (3 large aircraft visiting the airport) due to concerns about safety. The main concern of safety is the short runway of 1.9 km, which is too short for larger aircraft too land in certain conditions. Some days later however, this closure ruling was overruled by an appeal, stating that the safety considerations to be taken into account did not outweigh the economic ramifications that would ensue due to airport closure. So basically all the people in the jet died because of money and we just found out the exact numeric value that the Infraero, ANAC, government, Justice have considered "equals" human life.

The Airbus A320 is able to land on the runway of that length in dry weather conditions. Not in wet weather conditions. There are accounts of the Airbus failing to engage the reverse thrust, as it happened in Warsaw and some other event (mysterious) in France some time later. In this case with SP, it appears that the airplane had problems with the reversor since the Friday before (which, ironically, was the 13th). According to TAM, this was not prohibitive to still using the airplane. I'll leave this to airplane experts to decide whether this is correct or not. The pilot attempted to take off again, (but very likely due to aircraft logic was unable to). At least someone is looking at flight deck automation problems.

There seems to be a strong will to make money in Brazil and this focus is costing lives of other people. Rather than complying with all standards, assume the responsibility beyond the will to make money, some people are playing russian roulette with other people's lives.

If there is no consequence this time around and the "guilt" remains in the middle as it has been the case for other incidents.... Brazil is hopeless. It will mean there is no accountability for Brazil, no conscience, no responsibility, and not even authority or leadership. If that be the case, get out while you can! Before you become another statistic.

On another note... the PanAm games are there. Millions spent in the Maracanã stadium on some silly sport events when the people outside the stadium are living in atrocious conditions. I'm saying this not to say... let's NOT have the panams... I'm saying this because the money spent on having the games, with the full entourage and so on, seems a bit much. The positive thing is that even the poor living on the famous slums hill can see the fireworks going off in several rounds during the opening concert. It was amazing and beautiful. That should make them at least slightly happier...? Or am I safe in assuming it makes them quite mad to see how money is being wasted on fireworks that could have been used to improve poor health-care or impossible sanitary conditions, or ... maybe... like... ending drugs and violence in Rio?

Ignorance abound! One day or another... Brazil will have to face its consequences. Or rather... the people will.

(image above courtesy Duke Chargista).

Tuesday, July 17, 2007

On the meaning of meaning...

I think through my reasoning of the previous posts there is a certain scope to mathematically represent certain concepts of meaning and their relationships in a different way than NLP does at the moment.

The challenge is:
  • Natural language embodies meaning (semantics)
  • The embodiment of this meaning should be extracted and translated to a different representation, ideally mathematical
  • The interrelations between concepts should be clarified and also encoded into a mathematical representation
  • A document should be analyzed according to a world model or instance model that a large network may have. Then generate a representative network model of the meaning of that document within that world model or instance model
  • I make the distinction between what I call model meaning and what I call instance meaning in that model meaning is something that applies to all instances (the truth of an instance), whereas an instance may differ because it has different or additional concepts or elements that do not or not always apply to the model meaning. An instance is easily recognized in (correct) language by words as "he", "its", "his", "her", "them". Things that belong to someone or things/concepts that have a specific name or identifier. General concepts do not have these names or identifiers.
  • Encode a query into a network model translation and disambiguate if necessary. Then find all network model translations that have similarities to the key network model
A further challenge in this topic is that just storing a network model of the overall meaning of a document is not enough, because lookup of that document through its meaning requires additional computational effort.

The necessity is to encode that particular meaning into a different key, such that this key has a specific meaning or range of meanings with error that can be used to look up the pertaining document. It should work the same way as storing a word that is referenced to a range of documents. Knowing the word, we can look it up from the database and retrieve all documents in which the word occurred.

For meaning, this is obviously very different and far from straight-forward, plus that there is very likely a large margin of error in analyzing its meaning (use of synonyms adds to this error and might also slightly change the meaning if changed by a single choice of synonym).

It would be great to choose a very long number for example, which properly resembles the induced meaning of the document and where the document itself generates a range of different possible meanings that can be expressed as close to the generated number. This allows a query to be more effective and find a wider or smaller range of documents.

Monday, July 16, 2007

The problem: inferring meaning for computers

I browsed Wikipedia on the "Meaning of meaning". In order to allow computers to search the web semantically, it is necessary to allow a computer to understand meaning or at least map it to a category/number/element, so that it can infer relationships between words, passages and texts overall (between documents). I reckon this is computationally very intensive. It is necessary to better understand the concept of meaning in an attempt to represent it for a computer.

Well, reading Wikipedia, which is of course not the best reference on knowledge but acceptable for starters like me, I see that there are a number of very difficult problems arising when mapping meaning towards a mathematical element.

Meaning is induced by the environment and the interpretation of elements of a language. One text noted that knowledge is not stored as a linear corpus of text in the mind, but rather more like a network of elements that together represent the idea or concept. This means that rather than recalling the text corpus that describes the idea (after reading it the first time for example), knowledge is continuously reconstructed from the stored elements that we find (individually) important and relevant. This seems to mean that memory and the method how things are stored are very relevant for semantics. This explains also quite well how interpretation (based on experience) allows one person to totally misunderstand another, even though the language may be correct.

The problem with computers is that they are in general stateful (stacks, memory, CPU cache) and process one thing at a time. Consider for example the following paragraph from Wikipedia:

"In these situations "context" serves as the input, but the interpreted utterance also modifies the context, so it is also the output. Thus, the interpretation is necessarily dynamic".

It's easy to understand that when we process a certain corpus of text, the meaning and interpretation of that text will change as we scan it. This to me means that the analysis of a text in itself in one pass does not equate to the continuous, recursive analysis of that text, since the text itself is able to modify the context in which it is read. There is a feedback in the text that a computer will need to simulate. It seems that the more I read about semantics, the less I find computers able to simulate the mind processes that lead to understanding of meaning and communication of ideas. Let alone searching for it in a 400TB database (Internet).

Besides natural language in text form or speec, we are able to make sounds, facial expressions and we communicate through body language. The total of these elements will form a larger message that a computer cannot process. Also the emotional weight of certain texts is difficult to simulate for computers.

As I have written before, it does not seem possible at the moment to reliably construct a mathematical model for semantic search that works. There are only parts of the problem as a whole that can be simulated (a better word is approximated ).

Whereas it would certainly be very interesting to see whether semantics as a whole can be better approximated if we apply further matrix operations on matrixes of different purposes. For example, we could use LSI and LSA to consider relevance of one text to another on a very dry level, but multiply this with the knowledge of a particular context of reference, also represented in another matrix in the hope to find something more meaningful.

Matrices seem very useful in the context of deriving knowledge out of something we don't really understand :). A neural network is a matrix, LSI uses matrices and probably it's possible to come up with different matrices that represent contextual information or an approximation of context itself.

Assuming that we have a matrix for a concept or context, what happens when we apply an operation of that matrix on an LSI document? It may be far too early to do that however. In order to come up with anything useful it's necessary (from the perspective of the computer) to come up with a certain processing pipeline for semantic search.

These efforts probably also require us to re-think Human Computer interaction. A lot of our communication abilities are simply lost when we interact with a computer over the keyboard, unless we assume that our ability to communicate those concepts through language is very precise. As I said before, when we communicate and we communicate with people that have similar experiences, the level of detail in the communication need not be very large. This is because the knowledge reconstruction at the other end is happening more or less the same way (based on rather crude elements in the communication), which means that a lot of details are not present in the text. A computer might then find it very difficult to reconstruct the same meaning or apply it to the right/same context.

A further problem is the representation of knowledge, context and semantics. We invented data-structures like lists, arrays and trees that represent elements from quite restricted sets. The choice between these structures is governed by the general operation that is executed upon them and decisions are led by resource or processing limitations. However, the data structures were generally developed on the basis that the operations on them were known beforehand and the kind of operation (and utility of each element) is known at or before processing time.

Semantic networks (or representation of knowledge and/or context) do not exhibit this requirement, seemingly:
  • A representation of a concept, idea or element is never the root of things, or at least not a root that I can easily identify at the moment. Does the semantic network have a root at all? I imagine it more to be an infinitely connected network without a specific parent, a network of relationships.
  • The representation of a network in a computer data structure is not basic computer science.
  • Traversing this network is very costly.
  • The memory requirements for maintaining it in computer memory as well.
  • It is unclear how a computer can derive meaning from traversing the network, let alone apply meaning to the elements for which it is traversing the network.
  • Even if there are specific meanings that can be matched or inferred, the processing power is likely very high.
  • The stateful computer is not likely to be very helpful in this regard.
The latter is based on my imagination that the mind does not maintain a lot of state, but seems more a very rapid "functional language computer". Rather than retrieving meaning A or meaning B from memory directly based on the factors of a lookup, it reconstructs a meaning from smaller elements.

This goes back to a philosophical discussion on what the smallest elements of meaning are and how they interact together.

Latent Semantic Analysis

This is a wonderful explanation of LSA:

http://lsa.colorado.edu/whatis.html

"As a practical method for the statistical characterization of word usage, we know that LSA produces measures of word-word, word-passage and passage-passage relations that are reasonably well correlated with several human cognitive phenomena involving association or semantic similarity. Empirical evidence of this will be reviewed shortly. The correlation must be the result of the way peoples' representation of meaning is reflected in the word choice of writers, and/or vice-versa, that peoples' representations of meaning reflect the statistics of what they have read and heard. LSA allows us to approximate human judgments of overall meaning similarity, estimates of which often figure prominently in research on discourse processing. It is important to note from the start, however, that the similarity estimates derived by LSA are not simple contiguity frequencies or co-occurrence contingencies, but depend on a deeper statistical analysis (thus the term "Latent Semantic"), that is capable of correctly inferring relations beyond first order co-occurrence and, as a consequence, is often a very much better predictor of human meaning-based judgments and performance.

Of course, LSA, as currently practiced, induces its representations of the meaning of words and passages from analysis of text alone. None of its knowledge comes directly from perceptual information about the physical world, from instinct, or from experiential intercourse with bodily functions and feelings. Thus its representation of reality is bound to be somewhat sterile and bloodless."

Having read this from the perspective of inferring meaning from a corpus of text, I think perspectives and statements on the use of LSA or LSI are too positive to become anything truly useful for web search by itself alone.

A philosophical discussion on the meaning of meaning can be useful to understand how meaning is actually represented or can be analyzed. If ever we understand how meaning is derived, it should be possible to generate better approximate (mathematical?) models.

It's very difficult to infer any kind of meaning without having access to the real world the way that humans do. It would be interesting to find out how the world looks like to deaf or blind people. This should give us useful clues on the way a computer is perceiving a corpus of text. Moreover, maybe the way disabled people compensate can be a useful indication for other compensations in LSA or LSI.

It is very interesting though to see how meaning and semantics can be (in limited ways) represented by a mathematical calculation. This begs the question whether the mind itself is a large, very quick and efficient calculator or whether it's depending on certain natural processes. I think personally, as in another post, that the mind does not rely on calculation alone and that the model of a stack-based computer does not even come close to resembling our "internal CPU".

The intricate and complex process of deriving meaning from the environment requires an interaction between memory, interpretation, analysis and emotion. Mapping this to a computer:
  • Memory == RAM and disk, probably very, very large and not always accurately represented (human memory is 'fuzzy')
  • Analysis == Deconstruction of events into smaller parts
  • Interpretation == The idea inferred from the sum of the smaller parts, with extra information added from memory (similar cases)
  • Emotion == A lookup and induction of feelings based on the sum of the smaller parts, that recall certain emotions associated with the (sum of) those events. This is induced feelings when watching/reading a romantic love-story or in other cases levels of stress induced by a previously suffered trauma.
Clearly the computer is missing a lot of information. Besides the problems of Natural Language Processing (variations of meaning "hidden" in the text, where words mean different things, etc.), a poem to a computer is a sterile corpus of text that embodies much less meaning than it does to a human. Without memory and therefore association with similar events, a single corpus of text is empty and out of context.

These realizations lead me to believe that, in order for a semantic search to be really successful, one must replicate people's memories, emotions and contexts and analyze each corpus of text (the Internet) within the context of that particular person. To analyze and consider the whole Internet within the context of individuals is an impossible task. If we do this based on certain profiles, we might be able to execute this.

The ideal situation is the possibility to store "meaning" and not just keywords from a certain corpus of text and only later match this meaning with intention (search). I don't think we are able yet to represent meaning in other ways than text, unless we consider that LSA or LSI are indications of meaning by large arrays of numbers (matrices)?

Ugh! Sounds like LSD might be a better means to approximate meaning :)

Sunday, July 15, 2007

Latent Semantic Indexing

LSI (Latent Semantic Indexing) is a technique in computer science for finding certain "latent" information in documents. It's about analyzing semantic space through mathematics and statistics, which discovers semantic relationships between words and passages, however the computer cannot name that particular relationship. Also, actual meaning cannot be derived this way, but it can analyze how one corpus of text relates to another.

LSI creates a very, very large matrix of documents in columns with terms(words) in rows, where cells are occurrences. The to-be compared text is another single-column matrix that is transposed and multiplied with this very large matrix. The result is a couple of numbers that describe relevance, or similarity, both in the semantic space (not just word occurrence).

If you are interested, this tutorial gives a very good review of the technology. Several start-up companies are selling Search Engine Optimisation "solutions" based on LSI, but these are all mostly a fraud:

http://www.miislita.com/information-retrieval-tutorial/svd-lsi-tutorial-1-understanding.html

LSI is an attempt to discover "latent" information in documents in an attempt to make our search engine searches more useful. Semantic search is about searching for meaning, whereas most current search engines use word occurrence search (a very dry method of search). LSI by itself is far from sufficient to even approximate a true semantic search.

I have just played around with this technology using a couple of papers found through Google. LSI Tutorial:

http://www.miislita.com/information-retrieval-tutorial/svd-lsi-tutorial-4-lsi-how-to-calculations.html

The technology is computationally very intensive (well, since matrix operations are, and the set we are considering is, namely the Internet). If you wanted to use LSI properly, you'd have to index all documents on the Internet first, establish a matrix (that will never fit in memory) with the number of columns equal to the documents you have analyzed and the number of rows to the unique terms (words) you have encountered. Then establish a matrix with your search query that has as many rows as the other matrix. Then transpose and multiply. It's easy to see that this type of processing can't easily be done online for the volume of searches that are taking place.

The silent Violence

When you go to Brazil, you may here and there notice certain "customs" that are absolutely appalling and date waaaay back to a couple of centuries ago. I went to a restaurant and there was a "family" with enough money so that it could hire a baby-sitter. Not something like in Europe for the evening, but a baby-sitter for the whole week, including nights, so these people sleep in the home and often in the same room as the little kids. There are many families that can afford one, the baby-sitters are called "baba". Sadly though, these are also people that are most frequently entirely ignored, even though physically quite present (how can you miss the person that carries your kids, nephews or others around?).

Anyway, the other family comes in and their lovely little son is being carried outside by the baba. The little mongrel is crying and they enter the restaurant. The remark is not: "There is Sandra/Yvete/Lena/Marcia with João/Miguel/Rafael"... it is "There is João/Miguel/Rafael, but he is crying". The little mongrel gets a place inbetween the parents. The baba "finds" a place at the end of the table. Well, let's sketch this out:

So, we see the whole family together and the round circles are plates. The social divide in this case means that:




  • The baba was not talked to by anyone of the family, so sat there in a complete state of isolation
  • The baba was not offered a plate
  • The baba was not offered food to eat at the same table as the "boss"
  • The baba was sitting at a table that was physically separate from the others
  • All in the family thought it was perfectly normal
  • The baba looked pretty bored by all this isolation
I know people in London that have a baby-sitter, someone that is there during the whole week. She makes some good money though and she eats along with the people of the house if she wants to. She prefers to eat at other times when there are guests, but that is by her preference. When people are in the house, she is part of the people in the room, she's not required to move anywhere else, she sits on the same couch as the guests and she talks to other people, including the guests. However, in general, she follows where the children are. The guests, on arrival, do not ignore her and ask her how she has been doing and establish at least a small conversation.

So... I see a lot of silent violence in the way how certain people are treated here. More like slaves than people. On every perspective, social, human rights, labour rights, salary, these people are already not equal to their hosts. Then above all, you get treated as if you didn't exist. I would certainly not want to be ignored or treated this way, but luckily I do have a choice. Then people wonder where these levels of violence arise from... duh!

Other situations are for example some resort places around Brazil. We've been invited once to a meeting, where the guy in charge decided to reduce the salary by 20% just three weeks before Christmas, intending to "do good" to the "community", because it would reduce the costs of running the resort. Brazilian law prohibits decreases in salary. But do you think any worker is going to sue?

One documentary on housekeepers in Brazil shows that some are expected to work 10-12 hours a day. If you wanted to study to get into college, then that's tough! Some of these have children that they bring into work. But whenever the host's little mongrel brat starts crying for his toys or whatever, the housekeeper must give preference to the little brat over their own children. Worst of all, these ignorant bosses even believe that they are providing a very good opportunity for these housekeepers, because "without them" they wouldn't have food on the table.

For these dark-age families, you know the most appalling aspect of all... They expect their baba's and housekeepers to be 100% loyal and dedicated to them no matter the circumstances...

GAAAAHHH!

Luckily this is not the situation for all domestic workers and there are more and more "good bosses and families" around, but it is improving very slowly. You'll find some very happy domestic workers that are very dedicated, but that is because they are being treated properly, not ignored. Even the frequent guests treat them as part of the family.

Here's a nice opportunity for those individuals that hire domestic workers to make a good difference... and this is not something out-of-reach because it's being governed by a politician or governmental organization. It's directly within reach and every little change in perspective, behaviour or expectation makes enormous differences on the other side.

Friday, July 13, 2007

Semantic Intelligence

I'm reading up as much as I can about semantic search. What I find on the Internet so far are quite a number of marketing materials, which shows that the concept of semantics is still very new. The direction taken in these materials is generally the analysis of language, linguistics, attempting to re-create common sense in a computer, as if it were possible to allow it to reason.

I'm very skeptical about these approaches at the moment, but don't totally discard it. The problem with a computer is that it is a fairly linear device. Most programs today run by means of a stack, which is used to push information about current execution context. Basically, it's used to store contexts of previous actions temporarily, so that the CPU can perform other tasks either deeper or revert to previous contexts and continue from there.

I'm not sure whether in the future we're looking to change this computing concept significantly. A program is basically something that starts up and then, in general, proceeds deeper to process more specific actions, winds back, then process more specific actions of a different nature.

This concept also more or less holds for distributed computing, for many ways this is implemented today. If you look at Google's MapReduce for example, it reads input, processes that input and converts it to another representation, then stores the output of the process towards a more persistent medium, for example GFS.

I imagine a certain model in the next paragraphs, which is not an exact representation of the brain or how it works, but it serves to purpose to understand things better. Perhaps analogies can be made to specific parts of the brain later to explain this model.

I imagine that the brain and different kinds of processing work by signalling many nodes of a network at the same time, rather than choosing one path of execution. There are exceptionally complex rules for event routing and management and not necessarily will all events arrive, but each event may induce another node, which may become part of the storm of events until the brain reaches more or less a steady-state.

In this model, the events fire at the same time and very quickly resolve to a certain state that induce a certain thought (or memory?). Even though this sounds very random, there is one thing that gives these states meaning (in this model). It is the process of learning. The process where we remember what a certain state means, because we pull that particular similar state from memory and that state in another time or context induced a certain meaning. In this case, analogy is then pulling a more or less similar state from memory, analyzing the meaning again and comparing that with the actual context we are in at the moment. The final conclusion may be wrong, but in that case we have one more experience (or state) to store that allows us to better define the differences in the future.

So, in this model, I see that rather than processing a many linear functions for a result, it's as if networks of different purposes interact together to give us the context or semantics of a certain situation. I am not entirely sure yet whether this means thought or whether this is the combination of thought and feeling. Let's see if I can analyze the different components of this model:
  • Analysis
  • Interpretation
  • Memory
  • Instinct, feeling, emotion, fear, etc.
That is interesting.

Well, the difference that this model shows is that semantic analysis talks about generally accepted meaning rather than individual meaning. The generally accepted meaning can be resolved by voting or allowing people to indicate their association when a word is on screen. This seems totally wrong. If for example a recent event, like 9/11 occurs, and the screen shows "plane", most would type "airplane" and the meaning of that word will very quickly distort other possible meanings: a surface, an "astral" plane, geometric plane, compass plane, etc. Meaning by itself doesn't seem to bear any relationship with frequency.

If this holds true, then it means that as soon as any model that shapes semantic analysis in computers has any relationship with frequency, it means the model or implementation is flawed.

Wednesday, July 11, 2007

Back "home"...

Well, finally made it back "home". I am part of a group called IACE (Instituto Antonio Carlos Escobar) which is a group of volunteers in Recife that are concerned about the rising violence levels. It's a good idea to get acquainted with the group and subscribe and help campaigning.

The site that was passed on the list is this:

http://www.pebodycount.com.br/

It's portuguese, but one post attracted my attention. I will translate it here and you should be aware of these things before you travel to Brazil:

-----

Balanço da violência nos seis primeiros meses de 2007.
São 2447 homicídios este ano e 2301, no mesmo período do ano passado.
A média atual é de 13 assassinatos por dia no estado. No ano passado eram 12.
Maio e junho de 2007 tiveram 733 assassinatos. O mesmo período de 2006, 718.
Considerando os dados do PEbodycount sobre o mês de junho:

85 assassinatos no Recife.
110 nas demais cidades da RMR. 35 em Jaboatão.
52 na Zona da Mata. 12 em Timbaúba.
49 no Agreste. 8 em Caruaru.
43 no Sertão. 9 em Petrolina.
11 em local indeterminado.

Entre as vítimas, foram contabilizados 330 homens e 20 mulheres.
Cerca de 80% dos assassinatos foram cometidos com a utilização de arma de fogo. Não estamos aqui apenas para fazer cálculos. Por trás desses números está a nossa realidade, que infelizmente, é a traduzida por essas estatísticas como sendo muito violenta. Estamos trabalhando para fornecer um elemento vital para a cidadania: informação. Façam bom proveito.

-----

Balance of violence in the first six months of 2007. There are 2447 murders this year (up to june) and 2301 in the last year, same period. On average, there are 13 killings per day in this state. Last year the average was 12.

May and June 2007 there were 733 murders. The same period 2006, 718.
Considering the data of PEBodyCount in the month of June only:

85 murders Recife. (3 million people) 110 in other cities of Regio Metropolitana Recife (6 million people, poorer neighborhoods) 35 in Jaboatão. 52 in "Zona da Mata". 12 in Timbaúba. 49 in Agreste. 8 in Caruaru. 43 in Sertão. 9 in Petrolina. (you should be able to find these places on Google Maps). 11 of unknown locality.

Between victims, 330 men and 20 women. About 80% of muders were committed with the utilization of firearms. We are not here to do calculations. Behind these numbers is our reality, that unfortunately, is translated through these statistics as being very violent. We are working to supply a vital element for citizenship: information. Take advantage of it.

-----

Mind you, I lived here since 2004, but I sense bad changes for the worse over the past couple of months. Today for example, I saw in the news that a shopping centre got totally terrorized by youngsters when cinema tickets went half-price. It's basically a very large group of 13?-2x year olds. Lots of robberies, violent assaults, verbal abuse against shopkeepers, stealing, vandalism, you name it! The shopkeepers were forced to close down, since they were unable to do their work this way. Shouting everywhere, people throwing stuff, damage, vandalism. There were about 60 men to provide security, many of which were off-duty police officers.

The police rep was interviewed at the same day and they saw no particular reason for concern and reinforced that the shopping center was as safe as normal "aqui há segurança sim".

Overall, people have already commented that this state is becoming (or has become), based on statistics, more violent than the state of Rio de Janeiro. Rio actually being the city that is most known for drug trade and illegal fire-arms.
(see "Cidade de Deus" for example).

Tuesday, July 10, 2007

Dublin, accountability, overreaction, swarm theory

I'm in Dublin at the moment to see some people here. It's a great city to be, albeit a bit rainy :). We're in a four-star hotel near the canal. Great stay, good food and real Guinness.

Shopping here is great too. The cost of living here is very high considering other places in Europe, possibly more expensive than Paris. Restaurants and so on are about 1.5x more expensive than common restaurants in London. We're going back tonight to London to travel back 6am towards Recife through Lisbon. Then it's back to work from Thursday to get the project finished.

I've had some interesting discussions around London and Dublin, also regarding Brazil. The industry and market for Brazil seems to get an international interest. There are English investors looking at certain regions and the stock exchange is gaining a lot of interest which is very good for Brazilian companies. I hope this means a turn-around time for the economy in Brazil, as there is some catching-up to do in that regard. Violence however is very difficult to push down at the moment.

With regards to corruption and violence levels, these discussions were putting forward the hypothesis that this is mostly due to lack of accountability. This does not just mean having watchdog organizations in place that signal occurrences, but having watchdog organizations that have the power themselves to do something about it. It is in the public interest if a certain violator of social norms and ethics gets fried significantly through the watchdog and through the common press, which requires a free press plus the insight of the politicians in the country that an independent organization with power must exist in order to progress as a society.

Nothing is perfect however. If you consider the UK, some things that shouldn't happen still do, but it's at a much different scale. Holland and the Scandinavian countries seem to do really well. It's key to keep a cool head and consider the situation from different viewpoints before acting. In a different post, I referenced a post from Bruce Schneier with regards to rare risk and overreactions.

Well, it's very specific to terrorism, but you could extend this observation into a deeper analysis of the human psyche, thus psychology. In how far are we as human beings able to make rational decisions "all the time" that is in the best interest at that time?

There is some very interesting research being done on "swarm theory" as well, parts of which may be related to human beings. Although we have certainly a good amount of individuality and reasoning, by how far are all our decisions truly made "individual" and not subject to a certain sense of "swarm thinking"? I guess this is a very interesting part of research for a psychology department. And in this sense, it is even more interesting to see how a certain context for semantic search algorithms need to take this into account (or be able to disregard it even!).

Cool, for now... I am getting ready to leave this hotel and get back through the city towards the airport. I hope back in Recife the sun is shining!

Wednesday, July 04, 2007

Development with GWT, one year down the road

With GWT being relatively new, it is now approaching a final build and may very soon be upgraded to production status and taken out of Beta. This is a good time to reflect on one year of GWT usage and my experiences so far.

When GWT came out, I was on an overnight standby and had all the time to look into the technology. I looked at the samples and very quickly got very enthusiastic. Now after one year, I still see the immense value that the toolkit provides and since the first release, many very important and cool changes and additions were made.

When Java came out, there was a lot of hype around the technology. Nowadays the Java language is accompanied by good practices and some people even have written rather large books about Java Design and coding patterns.

There are things one should know for GWT development as well. Many of the regular J2EE patterns are not exactly applicable to GWT overall, so many of the historic J2EE patterns are simply not useful.

A couple of things that I think every project should think out before starting on a GWT project:
  • What is the final size of the project? Will it be possible to fit this in a single module (compiler memory usage) or is it necessary to modularize from the start? If so, how will this be modularized?
  • How to structure components, modules and libraries to facilitate re-use of development in other projects as well? Use of imports!
  • It is highly recommended to think about a strategy for history browsing and perma-links. See History class of GWT and "onHistoryChanged".
  • Focus on the re-use of widgets and make developers aware of the importance of abstraction and reuse. I have found it is much more important to develop components that can be reused in different contexts than it is to solve a particular problem at hand in a particular way. Make sure to review that code.
  • Test the application on different platforms and on different browsers as you go along.
  • Develop the application on different platforms too. It makes sense to use Linux with FireFox and Mozilla by 2-3 developers of the project, where the other half uses Windows with maybe different versions of IE.
  • If the project will get very large, consider running in "noserver" mode from the start of the project. You will have to model your development environment slightly differently to be able to achieve this.
  • Develop proto-types of screens from the start. Do not develop screens before the proto-types are ready and you have an idea of the final Look and Feel of the application overall.
  • Hire a CSS expert. Make sure that your CSS tag names and approaches are consistent, make sense with your developers and its development is aligned with the HTML approach that GWT embeds. Code the general look and feel into the "gwt-" tags. For cases that require different approaches, derive from other tags and apply the differences there.
  • Give back to the community those cool things you are developing or document your innovative approaches. Post on the GWT newsgroups or help out with the development of one of those UI libraries out there. Some companies require specific legal sign-offs for contributing to open source projects.
  • Program against interfaces where applicable, not against specific classes.
  • Usability becomes much more important on the Internet. It's not just putting together some HTML pages anymore. Make sure you have someone on the team that understands usability issues and can design useful, easy screens for people to use.
Hope that helps. Good luck in your GWT projects!

Wednesday, June 27, 2007

Global Warming (An Inconvenient Truth)

I watched the (now) famous film of Al Gore yesterday. It's about global warming and the lack of human response against CO2 emissions and our pollution. I recommend watching the film and establish your own opinion. When assessing the topic with other sources, please use credible sources and not believe every article that run-of-the-mill media throws at you. There is a lot of propaganda out there that is not based on scientific results, or attempt to derail scientific results so far that are clear indications. Watch out for that!

http://www.climatecrisis.net/

Some companies believe that complying with gas emissions and other more rigorous standards will put them out of business due to unfair competition. Well, maybe. But that would point out that our social standards, those standards that we are looking for as consumers, need a bit of a push, plus significant consensus-building at a UN-level to establish better environmental standards worldwide. Why bicker about trade equality and so on when in a couple of years time there won't be a planet to bicker on? It's quite gullible!

What if all of this is simply untrue as some magazines try to make us believe? Bruce Schneier wrote a very interesting article on rare risk and overreactions. I'm making this point, because the investment on fighting terrorism has been very high over the past few years due to immediate emotional involvement, while global warming is a much more serious problem, but not felt as immediately needing a resolution. So, in a way it applies to this article as well:

http://www.schneier.com/blog/archives/2007/05/rare_risk_and_o_1.html


So I highly recommend seeing Al Gore's film. I will change my ways and opinions because of it. While you're at it, watch this video too:

http://www.youtube.com/watch?v=5g8cmWZOX8Q

1992, Rio Earth Summit.... so.... what has actually happened since? Back 15 years ago, we had people that were already seriously concerned... Where are the significant changes and modifications that have taken place since then?

Even though we're adults and we think of ourselves as pretty smart in general, we ourselves still seem to act like children when it concerns the custody of this planet. This role of custodian seems to be necessarily replaced by the role of a full-time janitor. So, when do we start to clean up this mess and make sure it doesn't repeat itself?

Thursday, June 21, 2007

GWT 1.4, Tomcat 6 and Comet tutorial

I've been experimenting with Comet applications a little bit and integrated this with GWT 1.4. The tutorial is sort of working, but still needs a bit of work in order to become more stable.

The tutorial is here:

http://gtoonstra.googlepages.com/cometwithgwtandtomcat

Please send me any comments, bug fixes, etc..

Tuesday, June 19, 2007

Contextual search...

This blog is called "radialmind" for a reason. It is based on my perception that the mind is radial and not linear. The whole concept is rather easy to explain. It is easier to consume a book by going through the hierarchy of it, that is, the TOC, the individual chapters, the paragraphs and the lines than it is to read the book from start to finish. This is the same concept that people from "mindmapper" use for example to document your ideas. It's not a linear documentation, it is radial.

You'll notice that when you start to read the book linearly from the start, each time you hit a header or paragraph header, you need to tell your mind to switch context. That is, put the particular following content into a particular context. If the hierarchy of the book is very poor, it will be very difficult to follow and read. This is because, I believe, you need to start back at the core of your context and take a different path to another part of the context where you will fill in the information that you are going to read.

Searching the web has some similar problems. All major search engines produce linear search results. There have been some search engines that do this differently in a sort of "related-words" kind of way, but these have been very poor because it takes a long time to get to your actual context from the point where you are (or where the web page is, rather).

I think it makes sense for words to provide an initial search context and then connect to other contexts of through verbs.

This post ties back to my post about "semantic web search". Rather than focusing on "nouns", entities, we should focus on contextualizing information through verbs and interaction.

So... maybe... as a thought and discussion to develop... nouns provide an initial context to the search that may be dead wrong. But the verbs further contextualize your thoughts into the specific items that you are looking for?

Room for further thought in this particular domain...

Comet and Server push

There is something I'm checking out in my spare time. A technology that exists for a while called "comet". There are other alternatives in server-push, but I prefer first to look into the one that seems a bit more standardized and thought out. I've read other threads in newsgroups that state comet is overkill in many situations.

Server-push isn't actually "push" in the sense that the server initiates the connection. It's more like a delayed client-pull with features on the server that prevent excessive resource drainage.

In this model, the client connects to the server and waits for information. Browsers have timeouts, sometimes servers do, so it will reconnect every x seconds if there is no data sent over the link. So, on error it reconnects. If the connection errors x times, it will stop connecting and display an error about unavailable services.

Comet has in the meantime be implemented in Tomcat as well as Jetty. Jetty has documented in more details how they implement server-side processing and has some statistics about expected resource usage and server loads.

A problem with a server could be the number of connections, but the first thing that runs out are processing threads.

The comet model basically uses asynchronous I/O processing with thread pools. So it's a way to multiplex your client connections to get time allocation by one of the threads in the thread pool. This of course requires the same session and client information to be available. A thread gets assigned to a connection when there is data to be read or when an internal server event (through the application layer actually) writes data to a client channel.

Using this model, the browser will also have to implement some event switch. It receives the data and does something with it. This could be a chat window or popup and so on. In general, a browser can only have two connections to the server. This may be a little bit limiting for data transfers, since the comet connection continuously consumes one connection. This makes image loads and asynchronous data loads take longer and become serialized. One way around this is to use virtual hosting techniques to assign the comet connection to some kind of HTTP event server and the other connections to other web servers that process data requests. A potential problem here is data security with javascript applications that are not always allowed when connecting to a different server than where the script came from.

Well, the usual applications arise from comet technology. chat and so on. But there should be other possibilities as well if we get the network security right (and issues like NAT and so on!). Consider connections from browser to browser without an intervening server. The immediate services are basically event-driven client applications that react on server events. You might connect to a certain site and whatever is going on in the system will notify you of that occurrence. This is a very interesting feature for many sites, since the user will no longer be 100% responsible to pull all that information to him.

Moreover, from a database access point of view, perhaps persistence frameworks can finally become more consistent. If the server knows that you are watching some kind of information and possibly editing it, any other user that edits the information before you should cause an event to be sent to your browser to notify you there is new unseen data to consider. The browser might even retrieve the new data and show it alongside. It should not be too difficult to get this done. There are caching mechanisms for example that can help in detecting which objects are being viewed by whom and when objects get refreshed in the cache through an edit.

Saturday, June 16, 2007

High Precision Event Timer

I've been suffering a bit from a somewhat sluggish machine at times. Sometimes I run VMWare and that is especially annoyingly slow. I already posted the max_cstate thing (powersaving functionality).

Here is another thing I tried. "hpet=disable" on the kernel start line in /boot/grub/menu.lst.

What is HPET?

I am still experimenting with hpet disabled, but so far the Gnome desktop seems more responsive and things slightly faster. I used to get some 'stutters' on the mouse cursor sometimes and more processing times, but things seem to run better actually without the precision timer.

At the same time I noticed that some applications crashed suddenly. Not very frequent, but when load was high.

I'll keep this value for now and see what happens. I can always revert.

FaceBook. The new web?

Web 2.0 and YouTube gave us "user-generated content". It is where we post our videos, audio, photos, text, blogs etc. online for everyone to see. 90% of everything is junk (maybe like this blog :).

The other 10% is funny, interesting, insightful, challenging, or whatever. Some later developments are new ways to play around with that content or host even new things that people didn't think of before. There are a million ways for example that we can interact with one another. Yahoo Pipes is all about processing news and information and delivering it to you through a kind of processing pipe.

FaceBook
is slightly different. You can inject content and pictures on a simple level, but you can also host embedded applications integrated with FaceBook. FaceBook is a bit like an existing portal on the web somewhere and then you can request your services to be integrated through this portal and use their API to interact with other services of FaceBook. If you consider "infra-structure", this is what FaceBook provides. You provide immediate business logic that is hopefully new to everyone.

Here are examples of this new kind of thing. The previos link shows the reasoning behind FaceBook, which sounds very interesting.

One of the last lines reads:
the Facebook Platform is primarily for use by either big companies, or venture-backed startups with the funding and capability to handle the slightly insane scale requirements.
Yes. If something is really successful and with the current efficiency of our social networking capabilities, "novelties" travel through our network at an insane speed. Not necessarily faster than general broadcasting, but there's also no filtering by a third party in the case of broadcasters. It could be that a 3rd party through other interests decides to downplay or diminish a certain event, which, when taken as "raw information" might be very important for everyone to know.

These snowball effects can increase load on any server farm in an instant. If you manage to get your company's link on CNN, BBC or Slashdot or any other large site, you'll certainly be sure of a lot of traffic instantly that may last for a day or two. If you consider social networking sites where people might actually return daily, if the services provided there are really good there is an exponential growth pattern and insane growth requirements. Just ordered that big iron? The next day you'll order 10 more. Whoops, your bandwidth is running out. Whoops, the firewall got attacked. One angry user just launched a bot-net attack on your servers.

Infrastructure, infrastructure, infrastructure and lots of investment, instantly. And on the business side you need to keep things interesting, or the network will quickly drain out. What happens when another site comes up that offers similar services and something new that you didn't think off? Is there any sense of "loyalty"? You're not talking to individuals necessarily. It would be interesting to see how individuals behave as part of a social networking site. Do they exhibit more a kind of "flock" behaviour (they go where "the rest" goes?) or are their actions still based on individual decisions?

If we can recognize "flocking behaviour", this may be good when the business grows... but wow, it can be very bad for business if the flock heads the other way.. there is no stopping it!

Here is another interesting post on one of the facebook blogs:
There is a valuable lesson in all of this. There is a ton of money in developing platforms that make it easier for people to express themselves quickly and easily. Following this thread I can imagine the future value of virtual worlds such as second life where users can pick and choose everything down to their clothing, height, etc with the click of a button. Life is a story. Those applications (software as well as physical devices) that make it easier for people to share their story for others to watch unfold will be the ultimate winners when all is said and done.

Thursday, June 14, 2007

Semantic Search

I'm not an expert at websearch, but here goes... Some rants and ramblings on semantic web search.

I watched a program this week with a well-known philosopher. The program was about technology and media mostly, as well as social networking sites and so on.

One question asked during this program was whether semantic websearch would soon be a possibility and when exactly this is likely to be happening. The response was, from the philosopher, that he didn't think semantic search would ever take off and is basically dead in the water. The argument was that the context and meaning of certain words differs from one person to the next.

Although this is true, then maybe semantic search does not really mean searching for things in a general context that is known to be true, but search in specific contexts that the search engine understands belongs to that person, his perceptions and beliefs (formed by life experiences, human contact, environment, country culture, tradition and so on).

One thing that I suspect is not mostly used in web search is the verb. Most searches strictly use nouns, but the context of that noun can differ enormously if it is not accompanied with a verb. The verb would put things into a more specific context to a great amount, but it is not yet in a personalized context.

Steve Yegge blogs about the differences in "verb" and "noun" thinking from the perspective of a programming language. You could say that programming languages are in a way means of communication with a machine, to express ideas and so on.

Anyway, as I said, I have no idea to what amount search engines currently use verbs or contextualize searches to be more specific. It might consider search history as one way of improving hits, but this is not very reliable as our priorities and contexts can change very rapidly.

Regarding implementations of such a search engine... It would be a search engine that exists today with the added difference that user interaction (with user profiling) would add a context indication to particular pages. I don't think it is necessary to actually define all contexts prior to classifications. If you work with neural networks for example, the computer has no idea what it is doing, but the end result of each calculation comes close to what is expected.

It would be a great idea for research. To tie a neural network at both ends for a search engine and see what comes out. The difficulty with this neural network is of course how to heuristically define numbers based on the page... Or rather, how to encode the content of the page in such a way that together with the input of words and the user profile, the end result will be a particular score.

Another approach is to focus more on the verbs and start counting occurrences and take that as a contextual factor.

Perhaps the most limiting thing in search is that the search itself is badly expressed with words? I have a certain contextual idea of things that I am looking for... What is the best way to tell a machine to go looking for that particular context? We could store user's profiles, focus on verbs and all of that, but what about location or approximate location?

Some search engines provide advanced searches and this may be very helpful in this regard. In order to get anywhere, I guess it makes sense to include psychologists and anthropologists in the discussion to understand thought, expression and context better. There may be ways to convert these things in different ways to gain a more meaningful communication dialogue with a machine.

People mostly consider semantic search to be : "teaching the machine". Punishing it when the results are not what you are looking for, rewarding it when it is exactly on the mark. But if the context differs from one person to the next, there is a never-ending cycle of punishment and the machine just gets confused. Some things that are in the same context for everybody will get very high search ranks. But searching should be more effective than that. It should also aim to expose the niches.

Monday, June 04, 2007

Python

Tonight decided to take a look at Python. I know a bit of Perl and did create a couple of scripts in this "language". More formal and stricter languages like Java, C# and Visual Basic seem easier to learn, but compared to Python generate a whole lot more code.

I was pleasantly surprised by the offered modules from Python and how little code it takes to accomplish something. The documentation is quite up to speed and it offers some quite ingenious unit testing capabilities. You just wrap it into the docs. It then becomes a test case plus an example for somebody else to use.

Now, I find Python slightly easier to use in comparison to Perl. It's slightly less cryptic and uses more the concept of "function" than "operator". I was very pleasantly surprised to see it has support for SMTP, WWW, XML, UNIX, threads, concurrency, data types, (easy) iterators, functions, modularization, serialization, profiling, C module extensions, classes, embedding and so on. It's in a way sort of comparable with certain features of Java.

My advice: Give the tutorial a spin once... You'll get to know the capabilities and the rest from there is just reference work!

http://www.python.org/doc/

Low-level == innovation, high-level == entrepreneurship

If you work in a place where innovation is stimulated (like C.E.S.A.R), there are certain ideas that you can truly identify as innovation and others that are new or providing services that do not yet exist, but are not necessarily "innovation".

I'm looking at a couple of ideas all around for new technological "break-throughs" and a good number of these ideas, described as innovation, actually are "new systems" that just do not yet exist in the market. There are big problems here for the execution of these ideas:
  • Large systems are very risky to build and the effort required to actually complete them is always at least two times higher than the estimated effort. (see Vista, see internal "large" projects).
  • To make money of these "larger" systems is difficult. You start from a couple of clients maybe (if you succeed in selling it), but this is not the beginning when you start to make money. That is only after x years.
  • Support, after-care etc. are difficult to arrange. Your "development" doesn't just stop right after development is terminated. You may actually need more people to be able to seel than you did during development. Do you really want to start a company?
  • Maybe the system does not exist in the market, but is the problem actually more or less resolved in other ways? (is there truly market need?)
  • By the time the system is finished, the market may be gone.
Thus, you should consider:
  • If you want to start a business, aim for some niche market, develop a large system and focus on customers and selling product/services. This is not innovation in its entirety. Maybe only one line of code will be.
  • If you want to innovate, but not start a business as usual, focus on smaller parts at a lower level in the system.
Lower level innovations are for example search algorithms, distribution libraries, audio/video codecs, operating systems, embedded software, graph algorithms or the combination of the above to solve a *very* specific problem in technology. The focus should be on developing *technology* not directly applicable to any market problem. You'll have to focus on better efficiency in most of the cases of innovation.

Monday, May 28, 2007

Congress in Porto Alegre

I've come back from Porto Alegre, where I visited a congress on business administration (export, import and business between France and Brazil). There were some good presentations and the overall event was interesting.

We actually know a couple of people from this region, my brother in law and another couple that spent some time in Cambridge, UK to work with specific professors on their topic of interest. We met with them also. It was a lot colder than Recife.

Just on the outskirts of the city is the lake and a region they call "Ipanema". I had a look at the larger park close-by, which was used to host the "International Social Forum" some years ago.

Well, there was a bit of time left in the evenings luckily, which we spent with my brother-in-law to visit some good restaurants in the area. This particular area had a good lot of serious italian restaurants that I can recommend. Another one was the "Bistro da Rua", which is also a good place to just drink a bottle of wine and chat with some friends.

Tuesday, May 22, 2007

Linux Ubuntu on Dell

Dell is going to offer some hardware that comes with Ubuntu Feisty (7.04) pre-installed. These are exciting prospects for the propagation of Linux into mainstream markets that are non-geek. Does this mean that through the evaluation of Dell, Linux is finally considered mainstream-compatible?

Well, I know for one that Dell runs a lot of Windows software, not just offering it with their hardware, but they also run this internally. So they are not particularly interested in software as a technology (or as religion). We need to consider this move from Dell also from the perspective of support.

Windows, as the OS only and a couple of "productive" applications is a basic platform. Let's see the following formula:

known hardware + known software == x probable support calls

However, when we start installing "3rd party" software into this mix, suddenly the support calls can theoretically grow significantly larger. Not linear, but supposedly also exponential, especially when a commonly used piece of software happens to be incompatible with their hardware.

Ubuntu, as we know, has a repository with a very large amount of software on it that should satisfy most people's needs. I cannot truly remember the last time I downloaded an RPM or DEB from the Internet and installed it outside any repository.

This is a different kind of support that they have probably investigated, the support of testing the known software in the repository on their known hardware. If this provides all software tools that people need and it all works, then this is a fantastic service and may bring costs down significantly on the customer support front. Plus... since more than Dell customers use the software as well, for Dell support to resolve issues, they are no longer alone. They can count on the large number of forums, blogs and the community to help people out.

The question is thus... will Dell treat the community well? If they manage to become a responsible member of the community and contribute their productivty back, they gain much more than what they originally invested, as the momentum of the Ubuntu community should definitely increase. Openness in communication, testing of software on many different machines with many different kinds of users, hardware that works, software that works... Will this generate happy customers? Is this the future of computing? If mainstream picks up, does this mean that the market will *demand* Linux on their computers for community support (no fee!) and pre-tested distributions and hardware/software packages?

Some people are already discussing that Dell might start their own Ubuntu repositories and become full mirrors. If the cost of running those mirrors and hardware/software compatibility testing is lower than running a whole array of customer support reps, then this would actually lower the opex for Dell, which means more profit! exciting indeed...

Friday, May 18, 2007

Swamped...

Bit swamped at the moment, so little time to blog about anything. Will deploy a new release of Project Dune over the course of the weekend, which should include the new timesheet module. That will only be sort of Alpha-ish, since the locking functionalities won't work yet. Also, it needs a breakout class for the timesheet processor, so you can export your times reported to another system.

Well, great. We've been playing with GWT 1.4 for now. The hosted browser fixes and memory leak fixes really do help development, so you could say it really is maturing.

I'm looking forward to the suggestbox, number and dateformat classes and the Rich text editor for the document management part. Haven't had any time whatsoever to deal with HTML<->DocBook conversions though, let alone putting in place some simple templates for PDF conversions using XSL:FO.

Wednesday, May 09, 2007

Performance of VMWare on Linux disappointing?

Probably only for laptop with power-save functionalities... but you never know:

cat /sys/module/processor/parameters/max_cstate

echo 3 > /sys/module/processor/parameters/max_cstate

This value is related to the number of level that the processor takes to save energy (reducing its performance and power consumption when idle).

Friday, May 04, 2007

Computing challenges of the future

Since the first computer we've come quite a long way in integrating the computer into our daily lives. This has made significant changes to the formation of our society, in the context of expectations, beliefs and ethics.

There will be a time though that all processes (activities) that we perform are already automated and then we'll have a number of companies competing on the same level. Consider an application. It performs a certain kind of activity for the user. We have applications that work on private data at home (photo editing, accounting, office work like writing a letter, and so forth). At the moment we see these activities becoming more "online", so that we can do this from any place and not just from home.

The majority of applications has "options" and functionalities available that the user can choose from. By using these options and functionalities, we come close to executing that particular activity that we really wanted to do. There are new activities that can be invented that were otherwise not possible, but in general it is a replication of another activity in a different context that is transferred to another context for another need. (like sending a letter not just by writing a card at home and then posting it, but sending it from your PC and now sending it from your phone, also called email). Different contexts, but essentially the same thing. Our society however changes its perceptions just by using that technology and more importantly, changes its expectations and constant needs. (people become dependent on the technology).

There will be a time though that many of these services or functionalities are already automated and that "aggregated value" by itself can not be provided by just automating such activity. It needs something else.

I guess that one of the challenges of the future is that the computers need to be more helpful in their assistance to making decisions or assistance in detection. This probably means a lot of data-warehousing and processing and analysis and then use the outcomes of this whole process. It goes way beyond any kind of simple algorithm in this case. We need to mathematically understand better how the world fits together.

So the future of IT seems more geared towards "optimization" and "efficiency" than just "automatization" the way it's still enjoying things. After we have automated the majority of tasks and all these companies offer their products, the next step is optimization of these tasks. To be able to better understand activities, you must understand more about that activity domain. Or rather, linking back to my previous post... a computer engineer will not be able to compete on higher levels any longer just by having knowledge about computing or engineering. You'll have to invest in gaining knowledge of other problem domains and then apply the knowledge of one domain to the other, cross-over.

Saturday, April 28, 2007

The IT worker of the future

We all know that the US has a large number of IT companies or departments in mega-corporations and we also know that over the past couple of years, a large number of people were deemed "redundant" in their own country. Their jobs were off-shored to other countries like India, China or Brazil, mostly because the actual work they were performing is similarly available as a service elsewhere at a much better cost. With the advent of the Internet, getting the results of that work and the project team interaction is very easy.

What we all probably need to do is re-think the objectives of companies, as I wrote about in earlier posts. Whilst I am not going to even start to highlight particular companies in this posts, (it is irrelevant), the concept and context of a corporation and company responsibility should be evaluated. Is it meaningful that a company is only aiming to make money in this world? What if we, as a democracy, change the legislation to make them also aim for other objectives? Should we introduce protective laws against job off-shoring?

If the trend continues, then I am afraid all IT workers will need to start learning more than just IT. The technology is getting easier all the time or at least more accessible and documented. Just look on any search engine for some particular problem and it is likely that you get many possible solutions. People with lower wages in other countries have access to that very same information. How can you make a difference?

I believe that much of actual software that is written today will become more commoditized. Not to the level of steel, just more. If this happens, then it is no longer enough to just know about technology. In the future, you will very likely need complementary knowledge in order to keep your job in your own country. You must bridge the gap between technology and some other activity that is really valuable to the business. This kind of activity is much more difficult to execute than following a requirements specification. It is also an activity that can possibly produce a lot more value for your employer than if it would just acquire an existing solution.

The first thing to consider is that IT doesn't really matter. It is there to be used, but it's a bit like a truck. The truck itself provides no value just by its existence, but the way how it is used is meant to bring the cargo from A to B. So, the shipment of cargo from A to B is the actual thing that provides value, not the truck moving from A to B. The actual activity of building systems is meaningless, because you do it for a certain goal. That particular goal has meaning, not the system itself. The more you understand this distinction, the better and more efficient the systems you will build are.

I think that the IT people of the future can no longer sustain themselves in some richer countries unless they understand how they should improve themselves to better contribute to actual goals. This is possible by (better?) applying IT knowledge onto a better in-depth knowledge about the problem domain. Just building systems isn't going to cut it anymore. Focus on the problem, find out everything about the problem domain, find out how it really works, then shape your technical solution around that.

So, knowing about IT is still important in order to apply that knowledge to the problem. But it becomes much more important from a value perspective to understand the actual problem domains in which we are working. You should aim for being that person that writes the specification for producing something basically.

Thursday, April 19, 2007

Workflow systems

I'm looking at workflow systems for work and I'm very much enjoying it so far. The main package I'm looking at is called "jBPM". It's quite a large package to understand.

Workflow programming, to begin with, is different from actual programming in that it focuses on the process and on the execution of things, not on appearance or transformation of data. What I quite enjoy is that this statement means that you focus on the middle part of the construction of a system, rather than focus on the beginning (screens, user input) or the very end (database) or how some requirement fits in with your current system in an architectural way.

So, in order to successfully construct workflow-based systems, you first need to change your thinking about IT and how systems get constructed. It's not necessarily a very (pedantically) 'clean' way of programming (all data is stored in proper elements and there is compile-time dependency between all objects and all that), but it provides much more decoupling between functions, which on thus on the other hand improves reusability and componentization.

You should think of workflow as a series of automatic or manual actions that need to take place, easy to re-arrange and reusable. These actions are arranged in a diagram, which specifies which actor or pool of actors can execute those actions. The final part of the diagram also states whether these actions are executed in parallel or in series. To top things off... the whole process "instance", that particular order or that particular holiday request will get persisted into a database when it has no way to proceed immediately because it is waiting for external actions. You do not get these features "automatically" from any software language.

The latter word "language" is specifically mentioned, because you will start to mix your domain language (say Java or C#) with another "graph-based" language, in this case XML. So, the XML specifies authentication, actors, pools, the "flow" of the process, whereas the domain-based language helps to express componentized functionality (such as performing a credit-check, updating databases, gathering user input and so forth).

If you re-read the last paragraph, you notice how different workflow-based programming essentially is and how it could provide enormous value to IT systems if it is done right.

As an example, if you know Mantis or BugZilla, you know from the developer's or tester's point of view what the screen looks like for bug manipulation. Well, this screen is a single screen with all components in editable form, but from a workflow point of view, it should be constructed differently every time with only those components required for that particular step... For example, if you start to update or categorize the bug, then you do not need all the fields available in the screen to do that. When you have a process that dictates the need for a CCB comment, then you also have far too many elements on your screen to do specifically just that.

The point is that in general, many applications show everything on screen that you may not necessarily care about at that time and the user needs to compensate for the lack of guidance by understanding the process documented elsewhere in a very large document... Wouldn't it be great if we could just look at a screen and get our business process right and know what to do at the same time?

The other thing I notice is that many SRS documents are not necessarily focused on the process, but are focused on data. This shows me that there must be a huge opportunity in this area to provide consultancy on core business processes and merge this with IT specific knowledge.

Software is something that should be changeable in an easy form. Some people interpret this as "scriptable", but scripts are software as well and you wouldn't want to have production-editable fields that allow business people for example to modify those scripts at runtime. So there are only specific scenarios in which scripts actually add value, because if only your developers modify those scripts, why not write it in a strongly typed language instead?

Workflow based systems, due to their high decoupling factor and focus on process, might be slightly easier to change than J2EE, .NET, C++ or ASP systems. It matters from the perspective of flexibility and how you allow systems to grow with changing needs, because the needs *will* change.

Lastly, someone at JBoss mentioned how workflow systems are very much in their initial phase, comparing it to RDBMS systems. It's not very clear how they can be used effectively or in particular environments... what else we need for workflow systems to become very successful. The key part in improving these opportunities is to take the best of both worlds and merge this into something better. We may have to take a radical step into something else and see where it goes. I am just considering... wf based systems may be slightly slower due to use of XML and process persistence... but with a couple of machines and putting all processes on similar systems, what do we need for widespread deployment of this technology?

There must be things missing here, because not everyone is using it yet :).

Sunday, April 01, 2007

Yet more

Here are yet some more design works:

Some cartoon art


galaxy explosion kind of thing


gel/glass/liquid button graphics. Can contain any text or icon really.



1024x768 wallpaper for my computer.

all of the above taken from tutorials to learn more about designs. The last took most time, but also was more actual drawing to do. The last has a very interesting, unintentional detail, which is the base of the tree. I sculpted the tree out of an actual picture and then noticed the trunk looks like a foot.

So, probably I'm going to read up a bit more on design theory from here on. Things like logo construction, colors and those kind of theories.

I actually tried some cartoon characters for drawing, but that didn't end up in much yet ;).

More design work

These are two more designs I created with Wacom and some standard functionalities.
This is a modification of a fractal. I basically heavily embossed it and layered the result on top of the original. Then grayscaled part of the picture for a concrete effect and layered other images over it for a sort of glass like effect. As you can see, I don't master glass yet, which is why I started with the following image, one of an orb over a newspaper. That one looks much better.

This orb is produced solely under photoshop. It started with a simple sphere in a flat color, than I added an inner shadow effect to create some white at the bottom right (it should have more in the final image). Then another gradient in the middle, which was masked by a radial gradient. A hotspot in white at the top and then a selection of the initial orb that had a fill of black and transparency, which was scaled down to about 40% in height, slightly blurred and slightly moved to the right. The glasses and newspaper are from a photo. And the problem I run into obviously is not paying attention to the existing shadows in the picture :).

G>