Friday, July 20, 2007

Analysis of the crash...

I've been looking at the video now in step-mode. It's poor quality, but I can derive a couple of things that question the statement that the reverse thrusters had been in operation the whole time. Looking at the video and taking some screenshots however, I'm not so convinced this was the case.

(click on the pictures to see them larger)

Here are some considerations of mine that suggest a different sequence of events that seems to be a closer account:

http://www.youtube.com/watch?v=k6lO-eig_i0&mode=related&search=

Here we see the approach of the aircraft, right, where the plane sees the camera at an angle of about 300 degrees, head-up:


This image shows how the aircraft is passing the camera. The rain should be an indication of full-thrust working. The speed of the air at thrust surpasses the speed of the airplane by far. I see no waterspray being pushed in front of the airplane. Also, the size of the spray and its shape seem to indicate that the jet being pushed backwards is solely attributable to the wheels. There isn't even, as far as I can see, a buildup around the wings that indicate any kind of reverse-thrust to stop the plane at this point. Note that at this point the plane is probably around half the length of the runway. The Airbus A320 requires at least 300-500 meters of runway beyond the 1.9km that this runway has:


Here we see another camera recording the event. The plane is in view lower right and is just entering the camera frustrum:


Camera 10 has a detail view. Now we do see a waterspray being pushed forward even beyond the forward wheel. The buildup of spray around the body of the aircraft as a whole is noticeable. This is what you'd probably have to see in the first image (albeit the shape of the water spray would be longer due to higher speed), but at least the water should move higher and around the body of the aircraft if reverse-thrust is to be engaged. See that white cloud in this picture? I didn't see that in the previous pictures. Common sense tells me that if reverse thrust was working earlier, we should at least see a visible deceleration and similar upward-moving waterspray in the previous pictures.


Here we see another image a second later. In the overall movement of things, I don't see the aircraft noticeably slowing down, but it's as if it suddenly starts rolling out ( no thrust applied whatsoever ) and it's as if the spray of the reverse thrust is suddenly diminishing significantly. Has the pilot just decided to abort the landing at this point in an attempt to take off again with the remaining speed? This is probably about 3/4 down the runway or so:


This image here shows another camera about 3 seconds later. There seems to be a short flash at the left turbine, which may indicate the thruster reversing again into forward mode. I am not sure whether reversing the thrusters at this point, when still rotating reverse, would ignite a flash of some kind (maybe someone can comment?). The length of the runway in front must definitely be very limited.


Even though the news indicates that the right reversor was defective, I'm not sure whether this is truly the cause. It's very well possible that it was not operating and the aircraft logic prohibiting the operation of the left reverse thruster. However, there are other accounts of disabled reverse thrust or braking due to failure of the sensors or conditions prohibiting proper sensing of ground conditions (the A320 does not allow pilots to engage reverse-thrust in "flight" condition).

http://www.aaib.dft.gov.uk/publications/bulletins/february_2005/airbus_a320_200__c_ftdf.cfm

http://www.rvs.uni-bielefeld.de/publications/Incidents/DOCS/ComAndRep/Warsaw/warsaw-report.html

http://nakedshorts.typepad.com/nakedshorts/2005/08/debugging_airbu.html

Plus... we see that the reverse thrust does kick in at some point (picture 4), albeit much too late. If there was a full defect, this wouldn't have happened.

There seems to be quite some confusion with engineers as to how the actual braking operation works, as the manual is not very clear at this point:

http://www.pprune.org/forums/archive/index.php/t-92017.html

I can imagine that when one engine does not allow reverse thrust, the other should not apply it as this would spin the airplane around. There should be a safety mechanism in an aircraft that guarantees (more or less?) equal thrust being applied to both engines. Steering in an aircraft is not done through thrusters, but through wing action and ailerons.

Personally, if this is what happened, I can understand the stress of the pilot. There is only just enough length to land in dry weather conditions. The runway is wet. The pilot with 20 years of experience must have landed here before and know about the length of the runway. The pilot must have known about the pending maintenance action for the right turbine. The grooving has not yet been done (pending for September) due to "commercial pressure" to open the airport ( the losses would be too great ). The aircraft on touchdown does not respond to any braking commands (see potential similarity with other reports on other A320 crashes). More than halfway down the runway, the aircraft finally responds, but the length in front of the plane simply doesn't cut it. The pilots probably both decide to pull back up. The reverse-thrust is aborted and the thruster is set to forward thrust again. The speed of the aircraft at this point is far from favourable with only very little runway left. Is it possible that another 300-500m would have saved their lives? The A320 takes off and lands at around 160knots. This is 82m per second. The landing speed was above 160 knots at the time of touchdown.

Besides this plane having probably suffered a technical problem, we cannot ignore the other factors that have contributed to this. The pilots get informed that they better circle around for another landing attempt if they do not manage to land at the first 300m on the runway. Imagine the consequences if the braking system doesn't work and you're halfway down the runway (that is 8-10 seconds down). The decision you are forced to take in the next 10 seconds is crucial and every second is very, very crucial. Braking the airplane at more or less full-speed having only half of the runway left, a runway without grooves in wet conditions? This probably went through the pilot's minds at the last point of decision.

A couple of people in Brazil just put a value on the price of human life. For a jet full of people to crash in a busy airport, the monetary equivalent is the money that was made from February 2007 until now.

My expectation for any airport wherever in the world is that all international and safety norms are met, if not exceeded. THEN we can talk about weighing off extra security and safety measures versus economic benefit.

Thursday, July 19, 2007

Accountability (vs. "relaxa e morra")

I found a new, very insightful and interesting blog from Lucia Hippolito. Cientista política, historiadora e jornalista, especialista em eleições, partidos políticos e Estado brasileiro.

http://www.luciahippolito.globolog.com.br/

She also commented on the fact of lack of accountability across Brasil. With the following main observations through the text:
  • Accountability contains the idea that authority is a public servant. Elected or not, it has to be accountability for its actions to society.
  • Less stage and more debate, less uprisings and more interviews, less "law by ministry" (do other countries have this even?) and more attendance to the Congress.
Well... Accountability. Such a great word, and there is no portuguese translation! (ironic?)

When we look at the disaster of the airplane in São Paulo, some important "political processes" immediately kicked in. Nope. Not what you expect. Immediate investigations were ordered to try to blame it on the runway, but overall, the political world kept rather quite. The president has, unfortunately, not appeared on television in the last 72 hrs to send his condolences and show his commitment and compassion. Bit disappointing.

In February 2007, the airport was closed for 737, Fokker 100 (3 large aircraft visiting the airport) due to concerns about safety. The main concern of safety is the short runway of 1.9 km, which is too short for larger aircraft too land in certain conditions. Some days later however, this closure ruling was overruled by an appeal, stating that the safety considerations to be taken into account did not outweigh the economic ramifications that would ensue due to airport closure. So basically all the people in the jet died because of money and we just found out the exact numeric value that the Infraero, ANAC, government, Justice have considered "equals" human life.

The Airbus A320 is able to land on the runway of that length in dry weather conditions. Not in wet weather conditions. There are accounts of the Airbus failing to engage the reverse thrust, as it happened in Warsaw and some other event (mysterious) in France some time later. In this case with SP, it appears that the airplane had problems with the reversor since the Friday before (which, ironically, was the 13th). According to TAM, this was not prohibitive to still using the airplane. I'll leave this to airplane experts to decide whether this is correct or not. The pilot attempted to take off again, (but very likely due to aircraft logic was unable to). At least someone is looking at flight deck automation problems.

There seems to be a strong will to make money in Brazil and this focus is costing lives of other people. Rather than complying with all standards, assume the responsibility beyond the will to make money, some people are playing russian roulette with other people's lives.

If there is no consequence this time around and the "guilt" remains in the middle as it has been the case for other incidents.... Brazil is hopeless. It will mean there is no accountability for Brazil, no conscience, no responsibility, and not even authority or leadership. If that be the case, get out while you can! Before you become another statistic.

On another note... the PanAm games are there. Millions spent in the Maracanã stadium on some silly sport events when the people outside the stadium are living in atrocious conditions. I'm saying this not to say... let's NOT have the panams... I'm saying this because the money spent on having the games, with the full entourage and so on, seems a bit much. The positive thing is that even the poor living on the famous slums hill can see the fireworks going off in several rounds during the opening concert. It was amazing and beautiful. That should make them at least slightly happier...? Or am I safe in assuming it makes them quite mad to see how money is being wasted on fireworks that could have been used to improve poor health-care or impossible sanitary conditions, or ... maybe... like... ending drugs and violence in Rio?

Ignorance abound! One day or another... Brazil will have to face its consequences. Or rather... the people will.

(image above courtesy Duke Chargista).

Tuesday, July 17, 2007

On the meaning of meaning...

I think through my reasoning of the previous posts there is a certain scope to mathematically represent certain concepts of meaning and their relationships in a different way than NLP does at the moment.

The challenge is:
  • Natural language embodies meaning (semantics)
  • The embodiment of this meaning should be extracted and translated to a different representation, ideally mathematical
  • The interrelations between concepts should be clarified and also encoded into a mathematical representation
  • A document should be analyzed according to a world model or instance model that a large network may have. Then generate a representative network model of the meaning of that document within that world model or instance model
  • I make the distinction between what I call model meaning and what I call instance meaning in that model meaning is something that applies to all instances (the truth of an instance), whereas an instance may differ because it has different or additional concepts or elements that do not or not always apply to the model meaning. An instance is easily recognized in (correct) language by words as "he", "its", "his", "her", "them". Things that belong to someone or things/concepts that have a specific name or identifier. General concepts do not have these names or identifiers.
  • Encode a query into a network model translation and disambiguate if necessary. Then find all network model translations that have similarities to the key network model
A further challenge in this topic is that just storing a network model of the overall meaning of a document is not enough, because lookup of that document through its meaning requires additional computational effort.

The necessity is to encode that particular meaning into a different key, such that this key has a specific meaning or range of meanings with error that can be used to look up the pertaining document. It should work the same way as storing a word that is referenced to a range of documents. Knowing the word, we can look it up from the database and retrieve all documents in which the word occurred.

For meaning, this is obviously very different and far from straight-forward, plus that there is very likely a large margin of error in analyzing its meaning (use of synonyms adds to this error and might also slightly change the meaning if changed by a single choice of synonym).

It would be great to choose a very long number for example, which properly resembles the induced meaning of the document and where the document itself generates a range of different possible meanings that can be expressed as close to the generated number. This allows a query to be more effective and find a wider or smaller range of documents.

Monday, July 16, 2007

The problem: inferring meaning for computers

I browsed Wikipedia on the "Meaning of meaning". In order to allow computers to search the web semantically, it is necessary to allow a computer to understand meaning or at least map it to a category/number/element, so that it can infer relationships between words, passages and texts overall (between documents). I reckon this is computationally very intensive. It is necessary to better understand the concept of meaning in an attempt to represent it for a computer.

Well, reading Wikipedia, which is of course not the best reference on knowledge but acceptable for starters like me, I see that there are a number of very difficult problems arising when mapping meaning towards a mathematical element.

Meaning is induced by the environment and the interpretation of elements of a language. One text noted that knowledge is not stored as a linear corpus of text in the mind, but rather more like a network of elements that together represent the idea or concept. This means that rather than recalling the text corpus that describes the idea (after reading it the first time for example), knowledge is continuously reconstructed from the stored elements that we find (individually) important and relevant. This seems to mean that memory and the method how things are stored are very relevant for semantics. This explains also quite well how interpretation (based on experience) allows one person to totally misunderstand another, even though the language may be correct.

The problem with computers is that they are in general stateful (stacks, memory, CPU cache) and process one thing at a time. Consider for example the following paragraph from Wikipedia:

"In these situations "context" serves as the input, but the interpreted utterance also modifies the context, so it is also the output. Thus, the interpretation is necessarily dynamic".

It's easy to understand that when we process a certain corpus of text, the meaning and interpretation of that text will change as we scan it. This to me means that the analysis of a text in itself in one pass does not equate to the continuous, recursive analysis of that text, since the text itself is able to modify the context in which it is read. There is a feedback in the text that a computer will need to simulate. It seems that the more I read about semantics, the less I find computers able to simulate the mind processes that lead to understanding of meaning and communication of ideas. Let alone searching for it in a 400TB database (Internet).

Besides natural language in text form or speec, we are able to make sounds, facial expressions and we communicate through body language. The total of these elements will form a larger message that a computer cannot process. Also the emotional weight of certain texts is difficult to simulate for computers.

As I have written before, it does not seem possible at the moment to reliably construct a mathematical model for semantic search that works. There are only parts of the problem as a whole that can be simulated (a better word is approximated ).

Whereas it would certainly be very interesting to see whether semantics as a whole can be better approximated if we apply further matrix operations on matrixes of different purposes. For example, we could use LSI and LSA to consider relevance of one text to another on a very dry level, but multiply this with the knowledge of a particular context of reference, also represented in another matrix in the hope to find something more meaningful.

Matrices seem very useful in the context of deriving knowledge out of something we don't really understand :). A neural network is a matrix, LSI uses matrices and probably it's possible to come up with different matrices that represent contextual information or an approximation of context itself.

Assuming that we have a matrix for a concept or context, what happens when we apply an operation of that matrix on an LSI document? It may be far too early to do that however. In order to come up with anything useful it's necessary (from the perspective of the computer) to come up with a certain processing pipeline for semantic search.

These efforts probably also require us to re-think Human Computer interaction. A lot of our communication abilities are simply lost when we interact with a computer over the keyboard, unless we assume that our ability to communicate those concepts through language is very precise. As I said before, when we communicate and we communicate with people that have similar experiences, the level of detail in the communication need not be very large. This is because the knowledge reconstruction at the other end is happening more or less the same way (based on rather crude elements in the communication), which means that a lot of details are not present in the text. A computer might then find it very difficult to reconstruct the same meaning or apply it to the right/same context.

A further problem is the representation of knowledge, context and semantics. We invented data-structures like lists, arrays and trees that represent elements from quite restricted sets. The choice between these structures is governed by the general operation that is executed upon them and decisions are led by resource or processing limitations. However, the data structures were generally developed on the basis that the operations on them were known beforehand and the kind of operation (and utility of each element) is known at or before processing time.

Semantic networks (or representation of knowledge and/or context) do not exhibit this requirement, seemingly:
  • A representation of a concept, idea or element is never the root of things, or at least not a root that I can easily identify at the moment. Does the semantic network have a root at all? I imagine it more to be an infinitely connected network without a specific parent, a network of relationships.
  • The representation of a network in a computer data structure is not basic computer science.
  • Traversing this network is very costly.
  • The memory requirements for maintaining it in computer memory as well.
  • It is unclear how a computer can derive meaning from traversing the network, let alone apply meaning to the elements for which it is traversing the network.
  • Even if there are specific meanings that can be matched or inferred, the processing power is likely very high.
  • The stateful computer is not likely to be very helpful in this regard.
The latter is based on my imagination that the mind does not maintain a lot of state, but seems more a very rapid "functional language computer". Rather than retrieving meaning A or meaning B from memory directly based on the factors of a lookup, it reconstructs a meaning from smaller elements.

This goes back to a philosophical discussion on what the smallest elements of meaning are and how they interact together.

Latent Semantic Analysis

This is a wonderful explanation of LSA:

http://lsa.colorado.edu/whatis.html

"As a practical method for the statistical characterization of word usage, we know that LSA produces measures of word-word, word-passage and passage-passage relations that are reasonably well correlated with several human cognitive phenomena involving association or semantic similarity. Empirical evidence of this will be reviewed shortly. The correlation must be the result of the way peoples' representation of meaning is reflected in the word choice of writers, and/or vice-versa, that peoples' representations of meaning reflect the statistics of what they have read and heard. LSA allows us to approximate human judgments of overall meaning similarity, estimates of which often figure prominently in research on discourse processing. It is important to note from the start, however, that the similarity estimates derived by LSA are not simple contiguity frequencies or co-occurrence contingencies, but depend on a deeper statistical analysis (thus the term "Latent Semantic"), that is capable of correctly inferring relations beyond first order co-occurrence and, as a consequence, is often a very much better predictor of human meaning-based judgments and performance.

Of course, LSA, as currently practiced, induces its representations of the meaning of words and passages from analysis of text alone. None of its knowledge comes directly from perceptual information about the physical world, from instinct, or from experiential intercourse with bodily functions and feelings. Thus its representation of reality is bound to be somewhat sterile and bloodless."

Having read this from the perspective of inferring meaning from a corpus of text, I think perspectives and statements on the use of LSA or LSI are too positive to become anything truly useful for web search by itself alone.

A philosophical discussion on the meaning of meaning can be useful to understand how meaning is actually represented or can be analyzed. If ever we understand how meaning is derived, it should be possible to generate better approximate (mathematical?) models.

It's very difficult to infer any kind of meaning without having access to the real world the way that humans do. It would be interesting to find out how the world looks like to deaf or blind people. This should give us useful clues on the way a computer is perceiving a corpus of text. Moreover, maybe the way disabled people compensate can be a useful indication for other compensations in LSA or LSI.

It is very interesting though to see how meaning and semantics can be (in limited ways) represented by a mathematical calculation. This begs the question whether the mind itself is a large, very quick and efficient calculator or whether it's depending on certain natural processes. I think personally, as in another post, that the mind does not rely on calculation alone and that the model of a stack-based computer does not even come close to resembling our "internal CPU".

The intricate and complex process of deriving meaning from the environment requires an interaction between memory, interpretation, analysis and emotion. Mapping this to a computer:
  • Memory == RAM and disk, probably very, very large and not always accurately represented (human memory is 'fuzzy')
  • Analysis == Deconstruction of events into smaller parts
  • Interpretation == The idea inferred from the sum of the smaller parts, with extra information added from memory (similar cases)
  • Emotion == A lookup and induction of feelings based on the sum of the smaller parts, that recall certain emotions associated with the (sum of) those events. This is induced feelings when watching/reading a romantic love-story or in other cases levels of stress induced by a previously suffered trauma.
Clearly the computer is missing a lot of information. Besides the problems of Natural Language Processing (variations of meaning "hidden" in the text, where words mean different things, etc.), a poem to a computer is a sterile corpus of text that embodies much less meaning than it does to a human. Without memory and therefore association with similar events, a single corpus of text is empty and out of context.

These realizations lead me to believe that, in order for a semantic search to be really successful, one must replicate people's memories, emotions and contexts and analyze each corpus of text (the Internet) within the context of that particular person. To analyze and consider the whole Internet within the context of individuals is an impossible task. If we do this based on certain profiles, we might be able to execute this.

The ideal situation is the possibility to store "meaning" and not just keywords from a certain corpus of text and only later match this meaning with intention (search). I don't think we are able yet to represent meaning in other ways than text, unless we consider that LSA or LSI are indications of meaning by large arrays of numbers (matrices)?

Ugh! Sounds like LSD might be a better means to approximate meaning :)

Sunday, July 15, 2007

Latent Semantic Indexing

LSI (Latent Semantic Indexing) is a technique in computer science for finding certain "latent" information in documents. It's about analyzing semantic space through mathematics and statistics, which discovers semantic relationships between words and passages, however the computer cannot name that particular relationship. Also, actual meaning cannot be derived this way, but it can analyze how one corpus of text relates to another.

LSI creates a very, very large matrix of documents in columns with terms(words) in rows, where cells are occurrences. The to-be compared text is another single-column matrix that is transposed and multiplied with this very large matrix. The result is a couple of numbers that describe relevance, or similarity, both in the semantic space (not just word occurrence).

If you are interested, this tutorial gives a very good review of the technology. Several start-up companies are selling Search Engine Optimisation "solutions" based on LSI, but these are all mostly a fraud:

http://www.miislita.com/information-retrieval-tutorial/svd-lsi-tutorial-1-understanding.html

LSI is an attempt to discover "latent" information in documents in an attempt to make our search engine searches more useful. Semantic search is about searching for meaning, whereas most current search engines use word occurrence search (a very dry method of search). LSI by itself is far from sufficient to even approximate a true semantic search.

I have just played around with this technology using a couple of papers found through Google. LSI Tutorial:

http://www.miislita.com/information-retrieval-tutorial/svd-lsi-tutorial-4-lsi-how-to-calculations.html

The technology is computationally very intensive (well, since matrix operations are, and the set we are considering is, namely the Internet). If you wanted to use LSI properly, you'd have to index all documents on the Internet first, establish a matrix (that will never fit in memory) with the number of columns equal to the documents you have analyzed and the number of rows to the unique terms (words) you have encountered. Then establish a matrix with your search query that has as many rows as the other matrix. Then transpose and multiply. It's easy to see that this type of processing can't easily be done online for the volume of searches that are taking place.

The silent Violence

When you go to Brazil, you may here and there notice certain "customs" that are absolutely appalling and date waaaay back to a couple of centuries ago. I went to a restaurant and there was a "family" with enough money so that it could hire a baby-sitter. Not something like in Europe for the evening, but a baby-sitter for the whole week, including nights, so these people sleep in the home and often in the same room as the little kids. There are many families that can afford one, the baby-sitters are called "baba". Sadly though, these are also people that are most frequently entirely ignored, even though physically quite present (how can you miss the person that carries your kids, nephews or others around?).

Anyway, the other family comes in and their lovely little son is being carried outside by the baba. The little mongrel is crying and they enter the restaurant. The remark is not: "There is Sandra/Yvete/Lena/Marcia with João/Miguel/Rafael"... it is "There is João/Miguel/Rafael, but he is crying". The little mongrel gets a place inbetween the parents. The baba "finds" a place at the end of the table. Well, let's sketch this out:

So, we see the whole family together and the round circles are plates. The social divide in this case means that:




  • The baba was not talked to by anyone of the family, so sat there in a complete state of isolation
  • The baba was not offered a plate
  • The baba was not offered food to eat at the same table as the "boss"
  • The baba was sitting at a table that was physically separate from the others
  • All in the family thought it was perfectly normal
  • The baba looked pretty bored by all this isolation
I know people in London that have a baby-sitter, someone that is there during the whole week. She makes some good money though and she eats along with the people of the house if she wants to. She prefers to eat at other times when there are guests, but that is by her preference. When people are in the house, she is part of the people in the room, she's not required to move anywhere else, she sits on the same couch as the guests and she talks to other people, including the guests. However, in general, she follows where the children are. The guests, on arrival, do not ignore her and ask her how she has been doing and establish at least a small conversation.

So... I see a lot of silent violence in the way how certain people are treated here. More like slaves than people. On every perspective, social, human rights, labour rights, salary, these people are already not equal to their hosts. Then above all, you get treated as if you didn't exist. I would certainly not want to be ignored or treated this way, but luckily I do have a choice. Then people wonder where these levels of violence arise from... duh!

Other situations are for example some resort places around Brazil. We've been invited once to a meeting, where the guy in charge decided to reduce the salary by 20% just three weeks before Christmas, intending to "do good" to the "community", because it would reduce the costs of running the resort. Brazilian law prohibits decreases in salary. But do you think any worker is going to sue?

One documentary on housekeepers in Brazil shows that some are expected to work 10-12 hours a day. If you wanted to study to get into college, then that's tough! Some of these have children that they bring into work. But whenever the host's little mongrel brat starts crying for his toys or whatever, the housekeeper must give preference to the little brat over their own children. Worst of all, these ignorant bosses even believe that they are providing a very good opportunity for these housekeepers, because "without them" they wouldn't have food on the table.

For these dark-age families, you know the most appalling aspect of all... They expect their baba's and housekeepers to be 100% loyal and dedicated to them no matter the circumstances...

GAAAAHHH!

Luckily this is not the situation for all domestic workers and there are more and more "good bosses and families" around, but it is improving very slowly. You'll find some very happy domestic workers that are very dedicated, but that is because they are being treated properly, not ignored. Even the frequent guests treat them as part of the family.

Here's a nice opportunity for those individuals that hire domestic workers to make a good difference... and this is not something out-of-reach because it's being governed by a politician or governmental organization. It's directly within reach and every little change in perspective, behaviour or expectation makes enormous differences on the other side.

Friday, July 13, 2007

Semantic Intelligence

I'm reading up as much as I can about semantic search. What I find on the Internet so far are quite a number of marketing materials, which shows that the concept of semantics is still very new. The direction taken in these materials is generally the analysis of language, linguistics, attempting to re-create common sense in a computer, as if it were possible to allow it to reason.

I'm very skeptical about these approaches at the moment, but don't totally discard it. The problem with a computer is that it is a fairly linear device. Most programs today run by means of a stack, which is used to push information about current execution context. Basically, it's used to store contexts of previous actions temporarily, so that the CPU can perform other tasks either deeper or revert to previous contexts and continue from there.

I'm not sure whether in the future we're looking to change this computing concept significantly. A program is basically something that starts up and then, in general, proceeds deeper to process more specific actions, winds back, then process more specific actions of a different nature.

This concept also more or less holds for distributed computing, for many ways this is implemented today. If you look at Google's MapReduce for example, it reads input, processes that input and converts it to another representation, then stores the output of the process towards a more persistent medium, for example GFS.

I imagine a certain model in the next paragraphs, which is not an exact representation of the brain or how it works, but it serves to purpose to understand things better. Perhaps analogies can be made to specific parts of the brain later to explain this model.

I imagine that the brain and different kinds of processing work by signalling many nodes of a network at the same time, rather than choosing one path of execution. There are exceptionally complex rules for event routing and management and not necessarily will all events arrive, but each event may induce another node, which may become part of the storm of events until the brain reaches more or less a steady-state.

In this model, the events fire at the same time and very quickly resolve to a certain state that induce a certain thought (or memory?). Even though this sounds very random, there is one thing that gives these states meaning (in this model). It is the process of learning. The process where we remember what a certain state means, because we pull that particular similar state from memory and that state in another time or context induced a certain meaning. In this case, analogy is then pulling a more or less similar state from memory, analyzing the meaning again and comparing that with the actual context we are in at the moment. The final conclusion may be wrong, but in that case we have one more experience (or state) to store that allows us to better define the differences in the future.

So, in this model, I see that rather than processing a many linear functions for a result, it's as if networks of different purposes interact together to give us the context or semantics of a certain situation. I am not entirely sure yet whether this means thought or whether this is the combination of thought and feeling. Let's see if I can analyze the different components of this model:
  • Analysis
  • Interpretation
  • Memory
  • Instinct, feeling, emotion, fear, etc.
That is interesting.

Well, the difference that this model shows is that semantic analysis talks about generally accepted meaning rather than individual meaning. The generally accepted meaning can be resolved by voting or allowing people to indicate their association when a word is on screen. This seems totally wrong. If for example a recent event, like 9/11 occurs, and the screen shows "plane", most would type "airplane" and the meaning of that word will very quickly distort other possible meanings: a surface, an "astral" plane, geometric plane, compass plane, etc. Meaning by itself doesn't seem to bear any relationship with frequency.

If this holds true, then it means that as soon as any model that shapes semantic analysis in computers has any relationship with frequency, it means the model or implementation is flawed.

Wednesday, July 11, 2007

Back "home"...

Well, finally made it back "home". I am part of a group called IACE (Instituto Antonio Carlos Escobar) which is a group of volunteers in Recife that are concerned about the rising violence levels. It's a good idea to get acquainted with the group and subscribe and help campaigning.

The site that was passed on the list is this:

http://www.pebodycount.com.br/

It's portuguese, but one post attracted my attention. I will translate it here and you should be aware of these things before you travel to Brazil:

-----

Balanço da violência nos seis primeiros meses de 2007.
São 2447 homicídios este ano e 2301, no mesmo período do ano passado.
A média atual é de 13 assassinatos por dia no estado. No ano passado eram 12.
Maio e junho de 2007 tiveram 733 assassinatos. O mesmo período de 2006, 718.
Considerando os dados do PEbodycount sobre o mês de junho:

85 assassinatos no Recife.
110 nas demais cidades da RMR. 35 em Jaboatão.
52 na Zona da Mata. 12 em Timbaúba.
49 no Agreste. 8 em Caruaru.
43 no Sertão. 9 em Petrolina.
11 em local indeterminado.

Entre as vítimas, foram contabilizados 330 homens e 20 mulheres.
Cerca de 80% dos assassinatos foram cometidos com a utilização de arma de fogo. Não estamos aqui apenas para fazer cálculos. Por trás desses números está a nossa realidade, que infelizmente, é a traduzida por essas estatísticas como sendo muito violenta. Estamos trabalhando para fornecer um elemento vital para a cidadania: informação. Façam bom proveito.

-----

Balance of violence in the first six months of 2007. There are 2447 murders this year (up to june) and 2301 in the last year, same period. On average, there are 13 killings per day in this state. Last year the average was 12.

May and June 2007 there were 733 murders. The same period 2006, 718.
Considering the data of PEBodyCount in the month of June only:

85 murders Recife. (3 million people) 110 in other cities of Regio Metropolitana Recife (6 million people, poorer neighborhoods) 35 in Jaboatão. 52 in "Zona da Mata". 12 in Timbaúba. 49 in Agreste. 8 in Caruaru. 43 in Sertão. 9 in Petrolina. (you should be able to find these places on Google Maps). 11 of unknown locality.

Between victims, 330 men and 20 women. About 80% of muders were committed with the utilization of firearms. We are not here to do calculations. Behind these numbers is our reality, that unfortunately, is translated through these statistics as being very violent. We are working to supply a vital element for citizenship: information. Take advantage of it.

-----

Mind you, I lived here since 2004, but I sense bad changes for the worse over the past couple of months. Today for example, I saw in the news that a shopping centre got totally terrorized by youngsters when cinema tickets went half-price. It's basically a very large group of 13?-2x year olds. Lots of robberies, violent assaults, verbal abuse against shopkeepers, stealing, vandalism, you name it! The shopkeepers were forced to close down, since they were unable to do their work this way. Shouting everywhere, people throwing stuff, damage, vandalism. There were about 60 men to provide security, many of which were off-duty police officers.

The police rep was interviewed at the same day and they saw no particular reason for concern and reinforced that the shopping center was as safe as normal "aqui há segurança sim".

Overall, people have already commented that this state is becoming (or has become), based on statistics, more violent than the state of Rio de Janeiro. Rio actually being the city that is most known for drug trade and illegal fire-arms.
(see "Cidade de Deus" for example).

Tuesday, July 10, 2007

Dublin, accountability, overreaction, swarm theory

I'm in Dublin at the moment to see some people here. It's a great city to be, albeit a bit rainy :). We're in a four-star hotel near the canal. Great stay, good food and real Guinness.

Shopping here is great too. The cost of living here is very high considering other places in Europe, possibly more expensive than Paris. Restaurants and so on are about 1.5x more expensive than common restaurants in London. We're going back tonight to London to travel back 6am towards Recife through Lisbon. Then it's back to work from Thursday to get the project finished.

I've had some interesting discussions around London and Dublin, also regarding Brazil. The industry and market for Brazil seems to get an international interest. There are English investors looking at certain regions and the stock exchange is gaining a lot of interest which is very good for Brazilian companies. I hope this means a turn-around time for the economy in Brazil, as there is some catching-up to do in that regard. Violence however is very difficult to push down at the moment.

With regards to corruption and violence levels, these discussions were putting forward the hypothesis that this is mostly due to lack of accountability. This does not just mean having watchdog organizations in place that signal occurrences, but having watchdog organizations that have the power themselves to do something about it. It is in the public interest if a certain violator of social norms and ethics gets fried significantly through the watchdog and through the common press, which requires a free press plus the insight of the politicians in the country that an independent organization with power must exist in order to progress as a society.

Nothing is perfect however. If you consider the UK, some things that shouldn't happen still do, but it's at a much different scale. Holland and the Scandinavian countries seem to do really well. It's key to keep a cool head and consider the situation from different viewpoints before acting. In a different post, I referenced a post from Bruce Schneier with regards to rare risk and overreactions.

Well, it's very specific to terrorism, but you could extend this observation into a deeper analysis of the human psyche, thus psychology. In how far are we as human beings able to make rational decisions "all the time" that is in the best interest at that time?

There is some very interesting research being done on "swarm theory" as well, parts of which may be related to human beings. Although we have certainly a good amount of individuality and reasoning, by how far are all our decisions truly made "individual" and not subject to a certain sense of "swarm thinking"? I guess this is a very interesting part of research for a psychology department. And in this sense, it is even more interesting to see how a certain context for semantic search algorithms need to take this into account (or be able to disregard it even!).

Cool, for now... I am getting ready to leave this hotel and get back through the city towards the airport. I hope back in Recife the sun is shining!

Wednesday, July 04, 2007

Development with GWT, one year down the road

With GWT being relatively new, it is now approaching a final build and may very soon be upgraded to production status and taken out of Beta. This is a good time to reflect on one year of GWT usage and my experiences so far.

When GWT came out, I was on an overnight standby and had all the time to look into the technology. I looked at the samples and very quickly got very enthusiastic. Now after one year, I still see the immense value that the toolkit provides and since the first release, many very important and cool changes and additions were made.

When Java came out, there was a lot of hype around the technology. Nowadays the Java language is accompanied by good practices and some people even have written rather large books about Java Design and coding patterns.

There are things one should know for GWT development as well. Many of the regular J2EE patterns are not exactly applicable to GWT overall, so many of the historic J2EE patterns are simply not useful.

A couple of things that I think every project should think out before starting on a GWT project:
  • What is the final size of the project? Will it be possible to fit this in a single module (compiler memory usage) or is it necessary to modularize from the start? If so, how will this be modularized?
  • How to structure components, modules and libraries to facilitate re-use of development in other projects as well? Use of imports!
  • It is highly recommended to think about a strategy for history browsing and perma-links. See History class of GWT and "onHistoryChanged".
  • Focus on the re-use of widgets and make developers aware of the importance of abstraction and reuse. I have found it is much more important to develop components that can be reused in different contexts than it is to solve a particular problem at hand in a particular way. Make sure to review that code.
  • Test the application on different platforms and on different browsers as you go along.
  • Develop the application on different platforms too. It makes sense to use Linux with FireFox and Mozilla by 2-3 developers of the project, where the other half uses Windows with maybe different versions of IE.
  • If the project will get very large, consider running in "noserver" mode from the start of the project. You will have to model your development environment slightly differently to be able to achieve this.
  • Develop proto-types of screens from the start. Do not develop screens before the proto-types are ready and you have an idea of the final Look and Feel of the application overall.
  • Hire a CSS expert. Make sure that your CSS tag names and approaches are consistent, make sense with your developers and its development is aligned with the HTML approach that GWT embeds. Code the general look and feel into the "gwt-" tags. For cases that require different approaches, derive from other tags and apply the differences there.
  • Give back to the community those cool things you are developing or document your innovative approaches. Post on the GWT newsgroups or help out with the development of one of those UI libraries out there. Some companies require specific legal sign-offs for contributing to open source projects.
  • Program against interfaces where applicable, not against specific classes.
  • Usability becomes much more important on the Internet. It's not just putting together some HTML pages anymore. Make sure you have someone on the team that understands usability issues and can design useful, easy screens for people to use.
Hope that helps. Good luck in your GWT projects!

Wednesday, June 27, 2007

Global Warming (An Inconvenient Truth)

I watched the (now) famous film of Al Gore yesterday. It's about global warming and the lack of human response against CO2 emissions and our pollution. I recommend watching the film and establish your own opinion. When assessing the topic with other sources, please use credible sources and not believe every article that run-of-the-mill media throws at you. There is a lot of propaganda out there that is not based on scientific results, or attempt to derail scientific results so far that are clear indications. Watch out for that!

http://www.climatecrisis.net/

Some companies believe that complying with gas emissions and other more rigorous standards will put them out of business due to unfair competition. Well, maybe. But that would point out that our social standards, those standards that we are looking for as consumers, need a bit of a push, plus significant consensus-building at a UN-level to establish better environmental standards worldwide. Why bicker about trade equality and so on when in a couple of years time there won't be a planet to bicker on? It's quite gullible!

What if all of this is simply untrue as some magazines try to make us believe? Bruce Schneier wrote a very interesting article on rare risk and overreactions. I'm making this point, because the investment on fighting terrorism has been very high over the past few years due to immediate emotional involvement, while global warming is a much more serious problem, but not felt as immediately needing a resolution. So, in a way it applies to this article as well:

http://www.schneier.com/blog/archives/2007/05/rare_risk_and_o_1.html


So I highly recommend seeing Al Gore's film. I will change my ways and opinions because of it. While you're at it, watch this video too:

http://www.youtube.com/watch?v=5g8cmWZOX8Q

1992, Rio Earth Summit.... so.... what has actually happened since? Back 15 years ago, we had people that were already seriously concerned... Where are the significant changes and modifications that have taken place since then?

Even though we're adults and we think of ourselves as pretty smart in general, we ourselves still seem to act like children when it concerns the custody of this planet. This role of custodian seems to be necessarily replaced by the role of a full-time janitor. So, when do we start to clean up this mess and make sure it doesn't repeat itself?

Thursday, June 21, 2007

GWT 1.4, Tomcat 6 and Comet tutorial

I've been experimenting with Comet applications a little bit and integrated this with GWT 1.4. The tutorial is sort of working, but still needs a bit of work in order to become more stable.

The tutorial is here:

http://gtoonstra.googlepages.com/cometwithgwtandtomcat

Please send me any comments, bug fixes, etc..

Tuesday, June 19, 2007

Contextual search...

This blog is called "radialmind" for a reason. It is based on my perception that the mind is radial and not linear. The whole concept is rather easy to explain. It is easier to consume a book by going through the hierarchy of it, that is, the TOC, the individual chapters, the paragraphs and the lines than it is to read the book from start to finish. This is the same concept that people from "mindmapper" use for example to document your ideas. It's not a linear documentation, it is radial.

You'll notice that when you start to read the book linearly from the start, each time you hit a header or paragraph header, you need to tell your mind to switch context. That is, put the particular following content into a particular context. If the hierarchy of the book is very poor, it will be very difficult to follow and read. This is because, I believe, you need to start back at the core of your context and take a different path to another part of the context where you will fill in the information that you are going to read.

Searching the web has some similar problems. All major search engines produce linear search results. There have been some search engines that do this differently in a sort of "related-words" kind of way, but these have been very poor because it takes a long time to get to your actual context from the point where you are (or where the web page is, rather).

I think it makes sense for words to provide an initial search context and then connect to other contexts of through verbs.

This post ties back to my post about "semantic web search". Rather than focusing on "nouns", entities, we should focus on contextualizing information through verbs and interaction.

So... maybe... as a thought and discussion to develop... nouns provide an initial context to the search that may be dead wrong. But the verbs further contextualize your thoughts into the specific items that you are looking for?

Room for further thought in this particular domain...

Comet and Server push

There is something I'm checking out in my spare time. A technology that exists for a while called "comet". There are other alternatives in server-push, but I prefer first to look into the one that seems a bit more standardized and thought out. I've read other threads in newsgroups that state comet is overkill in many situations.

Server-push isn't actually "push" in the sense that the server initiates the connection. It's more like a delayed client-pull with features on the server that prevent excessive resource drainage.

In this model, the client connects to the server and waits for information. Browsers have timeouts, sometimes servers do, so it will reconnect every x seconds if there is no data sent over the link. So, on error it reconnects. If the connection errors x times, it will stop connecting and display an error about unavailable services.

Comet has in the meantime be implemented in Tomcat as well as Jetty. Jetty has documented in more details how they implement server-side processing and has some statistics about expected resource usage and server loads.

A problem with a server could be the number of connections, but the first thing that runs out are processing threads.

The comet model basically uses asynchronous I/O processing with thread pools. So it's a way to multiplex your client connections to get time allocation by one of the threads in the thread pool. This of course requires the same session and client information to be available. A thread gets assigned to a connection when there is data to be read or when an internal server event (through the application layer actually) writes data to a client channel.

Using this model, the browser will also have to implement some event switch. It receives the data and does something with it. This could be a chat window or popup and so on. In general, a browser can only have two connections to the server. This may be a little bit limiting for data transfers, since the comet connection continuously consumes one connection. This makes image loads and asynchronous data loads take longer and become serialized. One way around this is to use virtual hosting techniques to assign the comet connection to some kind of HTTP event server and the other connections to other web servers that process data requests. A potential problem here is data security with javascript applications that are not always allowed when connecting to a different server than where the script came from.

Well, the usual applications arise from comet technology. chat and so on. But there should be other possibilities as well if we get the network security right (and issues like NAT and so on!). Consider connections from browser to browser without an intervening server. The immediate services are basically event-driven client applications that react on server events. You might connect to a certain site and whatever is going on in the system will notify you of that occurrence. This is a very interesting feature for many sites, since the user will no longer be 100% responsible to pull all that information to him.

Moreover, from a database access point of view, perhaps persistence frameworks can finally become more consistent. If the server knows that you are watching some kind of information and possibly editing it, any other user that edits the information before you should cause an event to be sent to your browser to notify you there is new unseen data to consider. The browser might even retrieve the new data and show it alongside. It should not be too difficult to get this done. There are caching mechanisms for example that can help in detecting which objects are being viewed by whom and when objects get refreshed in the cache through an edit.

Saturday, June 16, 2007

High Precision Event Timer

I've been suffering a bit from a somewhat sluggish machine at times. Sometimes I run VMWare and that is especially annoyingly slow. I already posted the max_cstate thing (powersaving functionality).

Here is another thing I tried. "hpet=disable" on the kernel start line in /boot/grub/menu.lst.

What is HPET?

I am still experimenting with hpet disabled, but so far the Gnome desktop seems more responsive and things slightly faster. I used to get some 'stutters' on the mouse cursor sometimes and more processing times, but things seem to run better actually without the precision timer.

At the same time I noticed that some applications crashed suddenly. Not very frequent, but when load was high.

I'll keep this value for now and see what happens. I can always revert.

FaceBook. The new web?

Web 2.0 and YouTube gave us "user-generated content". It is where we post our videos, audio, photos, text, blogs etc. online for everyone to see. 90% of everything is junk (maybe like this blog :).

The other 10% is funny, interesting, insightful, challenging, or whatever. Some later developments are new ways to play around with that content or host even new things that people didn't think of before. There are a million ways for example that we can interact with one another. Yahoo Pipes is all about processing news and information and delivering it to you through a kind of processing pipe.

FaceBook
is slightly different. You can inject content and pictures on a simple level, but you can also host embedded applications integrated with FaceBook. FaceBook is a bit like an existing portal on the web somewhere and then you can request your services to be integrated through this portal and use their API to interact with other services of FaceBook. If you consider "infra-structure", this is what FaceBook provides. You provide immediate business logic that is hopefully new to everyone.

Here are examples of this new kind of thing. The previos link shows the reasoning behind FaceBook, which sounds very interesting.

One of the last lines reads:
the Facebook Platform is primarily for use by either big companies, or venture-backed startups with the funding and capability to handle the slightly insane scale requirements.
Yes. If something is really successful and with the current efficiency of our social networking capabilities, "novelties" travel through our network at an insane speed. Not necessarily faster than general broadcasting, but there's also no filtering by a third party in the case of broadcasters. It could be that a 3rd party through other interests decides to downplay or diminish a certain event, which, when taken as "raw information" might be very important for everyone to know.

These snowball effects can increase load on any server farm in an instant. If you manage to get your company's link on CNN, BBC or Slashdot or any other large site, you'll certainly be sure of a lot of traffic instantly that may last for a day or two. If you consider social networking sites where people might actually return daily, if the services provided there are really good there is an exponential growth pattern and insane growth requirements. Just ordered that big iron? The next day you'll order 10 more. Whoops, your bandwidth is running out. Whoops, the firewall got attacked. One angry user just launched a bot-net attack on your servers.

Infrastructure, infrastructure, infrastructure and lots of investment, instantly. And on the business side you need to keep things interesting, or the network will quickly drain out. What happens when another site comes up that offers similar services and something new that you didn't think off? Is there any sense of "loyalty"? You're not talking to individuals necessarily. It would be interesting to see how individuals behave as part of a social networking site. Do they exhibit more a kind of "flock" behaviour (they go where "the rest" goes?) or are their actions still based on individual decisions?

If we can recognize "flocking behaviour", this may be good when the business grows... but wow, it can be very bad for business if the flock heads the other way.. there is no stopping it!

Here is another interesting post on one of the facebook blogs:
There is a valuable lesson in all of this. There is a ton of money in developing platforms that make it easier for people to express themselves quickly and easily. Following this thread I can imagine the future value of virtual worlds such as second life where users can pick and choose everything down to their clothing, height, etc with the click of a button. Life is a story. Those applications (software as well as physical devices) that make it easier for people to share their story for others to watch unfold will be the ultimate winners when all is said and done.

Thursday, June 14, 2007

Semantic Search

I'm not an expert at websearch, but here goes... Some rants and ramblings on semantic web search.

I watched a program this week with a well-known philosopher. The program was about technology and media mostly, as well as social networking sites and so on.

One question asked during this program was whether semantic websearch would soon be a possibility and when exactly this is likely to be happening. The response was, from the philosopher, that he didn't think semantic search would ever take off and is basically dead in the water. The argument was that the context and meaning of certain words differs from one person to the next.

Although this is true, then maybe semantic search does not really mean searching for things in a general context that is known to be true, but search in specific contexts that the search engine understands belongs to that person, his perceptions and beliefs (formed by life experiences, human contact, environment, country culture, tradition and so on).

One thing that I suspect is not mostly used in web search is the verb. Most searches strictly use nouns, but the context of that noun can differ enormously if it is not accompanied with a verb. The verb would put things into a more specific context to a great amount, but it is not yet in a personalized context.

Steve Yegge blogs about the differences in "verb" and "noun" thinking from the perspective of a programming language. You could say that programming languages are in a way means of communication with a machine, to express ideas and so on.

Anyway, as I said, I have no idea to what amount search engines currently use verbs or contextualize searches to be more specific. It might consider search history as one way of improving hits, but this is not very reliable as our priorities and contexts can change very rapidly.

Regarding implementations of such a search engine... It would be a search engine that exists today with the added difference that user interaction (with user profiling) would add a context indication to particular pages. I don't think it is necessary to actually define all contexts prior to classifications. If you work with neural networks for example, the computer has no idea what it is doing, but the end result of each calculation comes close to what is expected.

It would be a great idea for research. To tie a neural network at both ends for a search engine and see what comes out. The difficulty with this neural network is of course how to heuristically define numbers based on the page... Or rather, how to encode the content of the page in such a way that together with the input of words and the user profile, the end result will be a particular score.

Another approach is to focus more on the verbs and start counting occurrences and take that as a contextual factor.

Perhaps the most limiting thing in search is that the search itself is badly expressed with words? I have a certain contextual idea of things that I am looking for... What is the best way to tell a machine to go looking for that particular context? We could store user's profiles, focus on verbs and all of that, but what about location or approximate location?

Some search engines provide advanced searches and this may be very helpful in this regard. In order to get anywhere, I guess it makes sense to include psychologists and anthropologists in the discussion to understand thought, expression and context better. There may be ways to convert these things in different ways to gain a more meaningful communication dialogue with a machine.

People mostly consider semantic search to be : "teaching the machine". Punishing it when the results are not what you are looking for, rewarding it when it is exactly on the mark. But if the context differs from one person to the next, there is a never-ending cycle of punishment and the machine just gets confused. Some things that are in the same context for everybody will get very high search ranks. But searching should be more effective than that. It should also aim to expose the niches.

Monday, June 04, 2007

Python

Tonight decided to take a look at Python. I know a bit of Perl and did create a couple of scripts in this "language". More formal and stricter languages like Java, C# and Visual Basic seem easier to learn, but compared to Python generate a whole lot more code.

I was pleasantly surprised by the offered modules from Python and how little code it takes to accomplish something. The documentation is quite up to speed and it offers some quite ingenious unit testing capabilities. You just wrap it into the docs. It then becomes a test case plus an example for somebody else to use.

Now, I find Python slightly easier to use in comparison to Perl. It's slightly less cryptic and uses more the concept of "function" than "operator". I was very pleasantly surprised to see it has support for SMTP, WWW, XML, UNIX, threads, concurrency, data types, (easy) iterators, functions, modularization, serialization, profiling, C module extensions, classes, embedding and so on. It's in a way sort of comparable with certain features of Java.

My advice: Give the tutorial a spin once... You'll get to know the capabilities and the rest from there is just reference work!

http://www.python.org/doc/

Low-level == innovation, high-level == entrepreneurship

If you work in a place where innovation is stimulated (like C.E.S.A.R), there are certain ideas that you can truly identify as innovation and others that are new or providing services that do not yet exist, but are not necessarily "innovation".

I'm looking at a couple of ideas all around for new technological "break-throughs" and a good number of these ideas, described as innovation, actually are "new systems" that just do not yet exist in the market. There are big problems here for the execution of these ideas:
  • Large systems are very risky to build and the effort required to actually complete them is always at least two times higher than the estimated effort. (see Vista, see internal "large" projects).
  • To make money of these "larger" systems is difficult. You start from a couple of clients maybe (if you succeed in selling it), but this is not the beginning when you start to make money. That is only after x years.
  • Support, after-care etc. are difficult to arrange. Your "development" doesn't just stop right after development is terminated. You may actually need more people to be able to seel than you did during development. Do you really want to start a company?
  • Maybe the system does not exist in the market, but is the problem actually more or less resolved in other ways? (is there truly market need?)
  • By the time the system is finished, the market may be gone.
Thus, you should consider:
  • If you want to start a business, aim for some niche market, develop a large system and focus on customers and selling product/services. This is not innovation in its entirety. Maybe only one line of code will be.
  • If you want to innovate, but not start a business as usual, focus on smaller parts at a lower level in the system.
Lower level innovations are for example search algorithms, distribution libraries, audio/video codecs, operating systems, embedded software, graph algorithms or the combination of the above to solve a *very* specific problem in technology. The focus should be on developing *technology* not directly applicable to any market problem. You'll have to focus on better efficiency in most of the cases of innovation.