Friday, November 2, 2018

The Kushner kid got a hotel out of the deal?

Iran sanctions are coming back. Here’s what you need to know

The Saudis bought a shit load of weapons and genocide is active in Yemen.  The Saudis are allied with our president, the Saudis, the same group that bombed us in 2001.  We should all be on strike.

Ten year at 3.2%

The pressure does not go away. We are in danger territory with Congress, we cannot afford a trillion per year in interest costs.

I got DLLs working

As in, load an attachment for join at run time, not at load time.

Took me a few hours to read the complicated instruction, but I eventually got it down to two basic conditions, declaring for export and load library function at run time.  I used CutnPaste from some the examples.

So I have a LoadAttachment function up and running, and it can be invoked from the command shell.  That completes all the connections I need, the system will not grow anymore, leaving all the spaghetti to the attachments.  I also now have a complete system able to test all the functions users will want.

The big news is I don't worry about cross platform, the load library has an equivalent in every OS. Someone else can put in configurations for all the different operating environments, letting me stick to the lab. Also, I do not need a system upgrade, gcc is working fine with its debugger and Notepad++.  I am not looking fopr IDE anymore, they have become bloated. I cannot use Virtual C IDE as I have already passed its capabilities.


Simple macro expander

This is backed by by a macro table of symbol expansions. If it finds the '$', then it immediately expands the macro right there in the argument list. Because it is recursive, it will expand any macros that are part of another macro. Works fin along with the symbol table, and good enough for building expressive control over cursors and matching inside join. It generates an argument list the same way c programs present argument list in the main routine, an array of arg pointers. So when I call a command, all the macros have already been expanded.

It can look for the assignment operator  '=', and output from a command is treated like a macro definition, placed in the table.  If their is no '$', it default to $Console, a built in macro.  Then $Console prints the result of the last command.  Very simple and good enough for the lab.

I have the freedom to develop commandlets for managing the join process knowing they can easily be transformed into standard shell macros. All the attachments, except the few core, will likely be DLLs, dynamically linked as needed.  I will write a simple Import commandlet that loads the dll file.


// Find arguments and expand macros
int ParseArgs(char ** ptrs, char * args) {
 int i=0,j;
 char *ch;
 Name * entry;
 do {
  while(*args && isspace(*args)) args++; // white space
  if(!*args) return(i);
  if(*args == '$') {
   args++;
   entry = get_entry(args);
   args=args+strlen(entry->name);
   ch = entry->expand;
   strmove(args,ch);  // move the string to make macro room
   j= ParseArgs(ptrs+i,args); // Drill down for any macros containing macros
   i = i+j; // Update the argument count
   return(i);
  }
  else {
   ptrs[i]=args;
   args++;
   i++;
  }

  while(*args &&  isalnum(*args) ) args++;
  if(*args) {
   *args=0;
   args++;
  } 
 } while(*args);
 return(i);  
}

Thursday, November 1, 2018

Simulating Power Shell

Getting my join to work as a Power Shell module requires me to 1) Upgrade, 2) Become more expert at Power Shell.

It is simpler for me to create a mickey mouse Power Shell-like macro function.  Then create more  expressive command functions for Match and Cursor in join. I know this is temporary, so I make it extremely simple that I can toss it at the right time.

The most likely outcome is that one of the tech companies or a start up runs with this idea, leaving me in the dust. The other possibility is that one of the companies makes a better Power Shell.

Example:

User may have four of five cursors engaged in a simple plain text file.   The user needs to group those cursor, and needs to change the match function, possibly on a Cursos bais. The user wants to create a batch file of the most interesting web pages so join can work through them And the shell needs a generic pipe function to express the concatenation of one join to another.

I can 'fake' a lot of this,write simplified version and have a jury rigged assignment function to simulate the variable assignment function in Power Shell.  Then I can move on with a competitive proof of concept the pros can integrate it into a shell.  Meanwhile, the package comes with enough shell to test all modes in a standalone environment without all the shell set up.

An argument parser is simpel,. so things like:

Create-Cursor arg1,arg2,... would create series of cursors attached to a particular attachment.
I can use
$result = Create-Cursor arg1,arg2,..

As long as I retain the same argument format, I can follow with:

Delete-Cursor $result  // Using the standard '$' for variable

This is how PowerShell works, it is easily duplicated in the special case for a single application.   Then the pros duplicate the idea, making it very integrated in PowerShell, keeping the join  concept.

So I did it

About 60 lines of code yields software that maintains macro expansions using '$' with the switch '-' character detected.  I Ijust did a crude implementation of the $ function inside a parse arguments routine, with absolutely no error checking and a strict syntax.

  But it is enough to simulate PowerShell; or any other shell dong macros. So I can that proceed to build cmdlets as needed to manage Cursors and Match functions while testing.  My code writing time was a couple of hours, but my system upgrade, set up, and learn process on Power Shell would have been two days, with uncertain results.

The power of c, do something with for proof of concept very quickly. And, as a bonus, we get a crappy little shell system usable until the market catches up with more powerful shells  Copy cat, the industry needs copy cats because it leads to a common standard for data.

More expressive control over join cursors and match functions

The typical AI willm have a dozen or more cursors joining various semantic networks. The user needs a more expressive method to group the cursors by category, and to swap out match functions in various attachments.  Requiring more thought, less coding.  We need a Power Shell for cursors and match functions, maybe even do this in a MS Power Shell, I dunno yet, I am on software break.

Huffman coded wordlists

What does -iLog(i) imply for a wordlist?

It is the probability the word list matches the text, we use match in a vector sense, not just on one atomic word.  Set the requirements for match, 3% of the text matches the list.  The goal is to adjust your word lists such that -iLog(i) applies then we have the structure, to a specified accuracy, we have the Huffman tree.

We can do this with ensembles of plain text, related. Works fine, but is more inaccurate for any individual selection of plain text.

Extend the -iLog(i) concept to obtain structure, the step and skip. You have maximum distillation, the amount of word list space needed by the individual searcher is minimized. It all boils down to the same thing, congestion on maximum entropy encoding trees, a much better optimization process than  neural nets.  Word lists requires much more training upfront, like Watson.  We have join technology for that.

HTML pages have structure, captured in the text scrape process, and in general,  most formatted text has structure easily found, we just need to write a scraper for each format. However the great deal of data ends up HTML  sooner or later, and we have that done.

Consider the case that word lists are visual feature sets, named with no assumed a priori  relationship. The the training consists of pruning word lists, again, u til you arrive at the optimum -iLog(i), and you have a finite but small guess as to the image..

One can see all of these techniques involve reducing identification as process of creating a balanced Huffman tree, all pruned lists have equal probability, uniform.  Then against the large set of randomly selected text, in the trained designated area of interest, the end user can try word list structures in sequence, equally probable.

Redneck Inc has competition

Google’s on-device text classification AI achieves 86.7% accuracy

Giggles wants to compete with my join system. I won. My system is simple, easy to convert to self learning.  

I was going to do taxes today, but now I have to read a bunch of technical articles about reducing sample size by 'distillation'. I don 't want Giggles to  whomp me.  My goal is to take their distillation technique and use it to narrow down word lists in classifying plain text.

What I do, let the pros move the work forward, then simplify their research.  Reading the stuff for the last five minutes makes me realize I need joint distributions from many word lists, and distill that down with a word list having the highest entropy (lowest redundancy) from the whole set of lists. Egad!

What is the Hamming distance between two words?


Image result for hamming distance
In information theory, the Hamming distance between two strings of equal length is the number of positions at which the corresponding symbols are different.
OK, we need equal length word lists. But we can be a bit flexible if we assume the large word list is more general and weight its distance accordingly.  But what exactly is the distance between two words?   We have a multi-dimensional problem, we can measure parts of speech, measure the root origination of words,. Measuring by prior classification works if the higher classification was trained on very large data sets, all of them of interest.

This is the problem at hand, data reduction across a reliable dimension when many dimensions are available.  The research is about how to do this with little prior knowledge, assume the prior classification is n on -parametric, we do not know the real distribution of any dimension.. Tough problem.  On the server side, join is free to stack word lists against training samples, and selectively refine the match process as we 'stack join's output to join's input.  The distance we really want are is available in the structure of the original text. Plain text can capture some of that rather than just generating a one dimensional list of plain words, some text processing is necessary, just enough to get the basic structure and we can find and narrow the wordlists at each node of our graph.  It is a mess of spaghetti, but doable.   I am still trying to get my brain around the concept, find the simple recursion, make it automatic..

In join we have the Match process, that is where we can subsample. I need a research staff, must go kidnap more Russian and Ukrainian mathematicians.

Think of  it as Huffman compression of sets of word lists.  Start  with the original plain text.  Match it against th4e general dictionary of commons word, larger than five characters.

 One can extract the run time distributions, the order and count at which words match. Do this in overlapping sections of the original text, generating multiple word list out.  The multiple lists out should contain the article structure, and using a Hamming distance between lists, reconstruct the step and skip of the original text.  That structural summary is given to the end user who can use it for more detailed and local searches on the text.

Anyway, that is the plan, me and a bunch of kidnapped mathematicians here at old Redneck Inc. be happy that we have stackable join technology.

Unconstitutional

Rein In the Administrative State -- and Preserve Democracy



By Peter J. Wallison missed US history. We are a republic, not a democracy.

Get your free housing in San Fraqncisco

San Francisco could spend nearly $700 million by the end of 2019 to support its homeless population if it adopts a proposal that would raise the money by levying taxes on big companies.The city is considering Proposition C, a proposal to increase taxes on large companies in order to improve conditions for homeless people living in San Francisco, The Wall Street Journal reported Thursday.Companies with over $50 million in revenue would be taxed “0.175 percent to 0.69 percent on gross receipts for businesses with over $50 million in gross annual receipts,” or “1.5 percent of payroll expenses for certain businesses with over $1 billion in gross annual receipts and administrative offices in San Francisco,” according to BallotPedia.If voters elect to adopt the ballot, the measure would produce an estimated $300 million in 2019 to be used for homeless aid — and San Francisco already spends $380 million in aid for the homeless population, according to TheWSJ.

California is our national refugee center, California and Texas.