Tuesday, September 11, 2012

Streaming

The natural tendency of both geography and biography is to stream things, to let nature take its course as diverse forces and pressures (forces confined, or spread out over areas) act on fluids. And so it is with human education.

The most effective thing is to let fluids find their level, and then channel them accordingly. The least effective thing is to do co-current heat transfer or to force all fluids to spread evenly in a thin and unnatural film over a thin and unnatural landscape (like some movies I've seen).

When people are streamed (at airport gates, by means testing, or in education), it is always a nett gain to society in terms of resources saved, as long as the streaming medium is relatively responsive to changes in the environment. What we shouldn't do is abandon any idea of streaming in the name of that soul-deadening anti-scientific anti-human thing called 'artificial equality'.

And that leads me to banding in school systems. If we no longer have it, is that a good thing? The point is that it is bad to avoid banding. This is because people do it anyway. It is biologically impossible for humans to avoid comparisons.

The better solution is to make ALL the data available and let humans test themselves against the data, making their own sense. Is a school that claims 'green' efficiency really so if it spends $500/head on electricity and water? Is it better to send your children to a school that is religious and cheap or secular and expensive? And what if one masquerades as the other?

What I propose is that far from scrapping banding, ranking and streaming, we should make all the data publicly available while suppressing the names of individuals. Show at least the bulk population data. For every school, release their accounts to the public for reasons of accountability. Let everyone see what schools do, so that people can make conscious pseudo-rational decisions that they can no longer blame on anyone else.

How's that for a different take? *grin*

Labels: , , , ,

Monday, November 07, 2011

Data Encoding

Today I had the misfortune to be looking at the latest draft output from one of the Ministries, those warded Grand Temples of the Meritokratia of Atlantis. The earnest young priests had coded two different functions with the same range of codes, then used that range of codes differently elsewhere.

Here is an hypothetical example of how it would have worked in practice.

Let's say that a person is the target of an imaginary ailment like Maladrach's Pyriferous Pustules. Then the possible procedures for treating it are MPP-01 (douse in water), MPP-02 (dowse for water), MPP-03 (pop the pustules), MPP-04 (pop the postulant), MPP-05 (postulate a pop). And the outcomes of the treatment might be MPP-06 (full recovery), MPP-07 (terminal spontaneous combustion), MPP-08 (referral to pyrothaumaturgic specialist) and so on.

Of course, you can't tell by inspection whether you're looking at a treatment or an outcome, since the code MPP is the same! Worse, if you have new treatments, you can't insert one after MPP-05, since you used running numbers for the outcomes that start immediately after MPP-05.

To make matters worse, these people decide that for smaller practitioners with fewer options, you could use MPP-01 (douse in water), MPP-02 (dowse for water), MPP-03 (full recovery), MPP-04 (referral to a specialist) and so on.

In other words, the same codes are used to represent completely different things in a different context.

If I had been coding, I would at the very least have used something more like MPPP-001 to MPPP-nnn for MPP Procedures, and MPPO-001 to MPPO-nnn for MPP Outcomes. And I would have reserved MPPP-nnnn and MPPO-nnnn for future use in case wonder-working turned out to be more complicated than the present state of the art.

It's a shame that the OSA probably prevents me from showing you, dear readers, the very slides themselves. Then you would see and be scandalized both by the great waste of your tax dollars and the calibre of the brains consuming said dollars.

Labels: , ,

Thursday, October 20, 2011

Responses 002 (Nov 2012)

The Nov 2012 list really throws up some interesting problems. Here is Question 2 on the list: "It is a capital mistake to theorise before one has data. Insensibly one begins to twist facts to suit theories, instead of theories to suit facts." — Arthur Conan Doyle. Consider the extent to which this statement may be true in two or more areas of knowledge.

(Note: As usual, the context of this statement is not supposed to be important. For the more conscientious amongst us, this quote comes from Conan Doyle's 1891 Sherlock Holmes story, A Scandal In Bohemia, in which Holmes is bested by the lady adventurer Irene Adler.)

The context which is important, however, is our consideration of the statement in terms of what a theory is supposed to be in different areas of knowledge. Here are a few thoughts.

In most conceptualisations of the scientific method, we're supposed to build theory from empirical data or reasoning from basic principles. It's either induction or deduction, repeated and mixed up, which generates theory. It isn't considered scientific to generate theory without any data at all, since you cannot even generate a problem statement or question that is of scientific value without some initial data. Why? Because a scientific claim must be testable, and we only know what a reasonable test is if we have some data to work with.

On the other hand, let's consider the arts. We'll have to define the arts as disciplines in which something material (a text, a narrative, an artifact etc) is created in order to induce a response (almost always emotional) based on somebody's sensory perceptions. A theory in the arts can be scientific in nature, if it is an analytical theory. However, a theory about what an artist thinks art is can seem spontaneously enacted through what he does. It's when people respond to that, that data are generated. The artist may not have any obvious foundations of data on which his theory of 'this is art' is built. He might be using his emotions as a guide, for example, or his intuitions, or his faith in his own arbitrary principles.

Other disciplines fall somewhere in between. All disciplines work with data; the question is whether data must precede theory (and thus be its foundation, as in the 'grounded theory' approach beloved of some research in the human sciences) or whether theory can be constructed before any data is received. It might be a chicken-and-egg kind of problem, requiring much thought before the obvious 'chicken came first' conclusion arises.

In this statement, however, is a lot more material for debate. You would have to think about 'capital mistake' (which implies 'fatal error'), 'insensibly' (by which Conan Doyle would have meant 'subconsciously', rather than 'irrationally', I think), and 'twist' (as in apply torque to deform, but not in a literal sense).

I suspect this question will be attempted by Sherlock Holmes fans, amateur investigators, and people who just want a no-holds-barred dust-up of the old-fashioned kind. A lot of fun, a lot of risk. "Two or more areas of knowledge," forsooth.

Labels: , , , ,

Saturday, July 02, 2011

The Cloud of Unknowing

The world is badly broken. The problem is a simple one: humanity has begun to assume that all difficulties can be overcome, all answers can be found, given enough cerebration and enough data. This is why we invest huge amounts of land, money, energy and hardware in storing enormous archives of data, most of which will be redundant, out of date, trivial, or impossible to convert into information.

The reality is sobering. We will be expending all of those resources because we think we should keep our stuff, our work, the unblessed fruit of our hands and minds, the debris of our data transactions and online interactions. We have overvalued the products of our thinking despite the fact that they have create problems we cannot solve, or created answers insufficient to those problems.

When London grew too crowded and too corrupt, the only solution for its chaotic and filthy tangle was the unthinkable Great Fire of 1666. Totally gutting the medieval Roman inner city, it destroyed the homes of seven-eighths of all London's inhabitants. It was this destruction of existing constructs that instructed the development of modern London and the architecture of Sir Christopher Wren.

Imagine this: what if ALL humanity's electronic data stores, in all forms and formats, media and mediating devices, were destroyed at one blow? A tragedy? A disaster?

I would say not. I would say it would be a golden opportunity to discover exactly what humanity is really made of, and to rebuild anew. However, I am not so sanguine as to think humanity will do better. Rather, I am inclined to believe that we will do worse.

Labels: , , , ,

Wednesday, April 29, 2009

Qualitative and Quantitative

Recently, I've had to speak with various students on the subject of what 'qualitative' and 'quantitative' mean, in the sense of qualitative or quantitative data or research methodology. It's interesting how oversimplication (the subject of my previous post) creeps in.

For instance, one of them said, "Quantitative means you can put a number on it."

I replied, "Well, you can put a number on a marathon runner and it doesn't make her quantitative data. You can have a school survey with ratings on items (a 'Likert scale') from 1 to 5, and that doesn't make it quantitative research either."

Quantitative data is actually data that can be gathered and placed within a scale of measurement, and can legitimately be subjected to mathematical operations. Sometimes this definition works to confuse people.

Take for example the idea of 'an average age'. Suppose that we ask your class how old they are in years (number of birthdays passed), and your class replies "17!" (40%) and "18!" (60%). Does this mean that the average age of the class is (0.4 x 17 + 0.6 x 18) = 17.6? That would make the average member of the class roughly 17 years, 7 months and 6 days old. What do you think? I suspect not. But the problem here is one of insufficient resolution and definition, not one of illegitimacy. Age can indeed be used as quantitative data in some contexts.

Qualitative data is actually data that specifies a kind or a class of property without being subject to scalar manipulation. It cannot be subjected legitimately to mathematical operation, although it is possible to try. This also confuses people.

Take for example the idea of 'colour'. Suppose we ask your class what their favourite colours are, and your class replies "Blue!" (50%) and "Gold!" (50%). Does this mean that the average favourite colour is metallic green? What do you think? I suspect not. The problem here is not numerical, but conceptual. You can't find an 'average' of favourite colours, since the average is unlikely to be anybody's favourite in this context.

The thing about this colour example is that it can be treated both quantitatively and qualitatively, with different kinds of results. A person can say, "My favourite colour can be expressed as that produced by a photon source in which all the radiation has the wavelength 530 nm. It's a kind of green."

Well, the wavelength is a manipulatable scalar, as is the intensity of the source and so on. But the greenness of the colour is what we call an example of the qualia, those sensory occurrences which we find difficult to think about, and which are almost by definition the basis of qualitative data. It is extremely unlikely that any two people will see exactly the same shade of green when exposed to a light source at 530 nm. This can be due to biology, biography or biasedness of some unknown kind. In fact, it is even less likely that the colour will affect them in exactly the same way.

This is why the methodologies that handle quantitative and qualitative data tend to be, respectively, mathematical and social. The former manipulates numbers of the kind that can be manipulated (vectors, scalars, cardinals — but not most ordinals); if any explanatory power resides in numbers, it is of the statistical and correlative kind. The latter tries to come to humanly acceptable consensus on what qualia could possibly have been observed and what kinds of explanation would be sufficient to account for them.

Of course, books have been written on guidelines for research in both kinds of methodologies and their deployment in many different areas of knowledge. I myself used mixed methodologies when doing my 1999 Master's thesis on Why Teachers Quit Teacher Development in Atlantis. Most of it was qualitative though; qualitative data is a lot better at explaining social phenomena and general insanity than quantitative data is. You can find my research online if you know where to look. Enjoy!

Labels: , ,

Sunday, April 20, 2008

Vectoring

The resultant is the sum of its constituent vectors; it is the statically equivalent result of what would happen if all vectors were applied individually to a body. More generally, the resultant is the final product of the application of a function to a set of data.

=====

What a statement that is! I have been researching the results of various schools and how they are reported in public documentation. I think that, quite often, the most commonly applied function is called 'spin'.

Essentially, the function spin transforms a set of data describing a negative trend into one that can look positive when viewed from the direction of the spin. This effect is not scalable; if the magnitude of the negative trend is very great, spin provides for very minimal difference. However, the effect of spin also magnifies the value of the largest positive trend. Since the function centres on such a value, it is entirely possible for spin to make things look better.

A careful look at spin shows that there must be some sort of critical point beyond which, no matter how much spin is applied, everything looks bad. A more careful look shows that since spin destroys information, applying an inverse spin function will not recover the original data. This means that you need to keep the original data if you want to see what existed before spin was applied. If you spin something enough, nobody will know how bad things really were.

Of course, it does not always suit some people to keep the original data. I have seen, in my time, the most outrageous misdirections being used so that bad results will appear good and good results will appear bad. Thank goodness I always keep the original data so that I can check my second-order (or greater) results.

Labels: , , ,