Showing posts with label high order functions. Show all posts
Showing posts with label high order functions. Show all posts

Tuesday, August 18, 2009

erlang: testing many conditions easily with lists of funs...

Sometimes you have to test many things before being able to choose the next action...

In many languages, you'll end up using a bunch of "if then else".
But in erlang, and the power of fun()s, you can efficiently write a simple function that will do all the job for you :p

Here is our purpose: call many functions with one argument.
For this example, we need to determine a file type with its filename.

Let's say that:
- the filename could be a valid 'word' temporary file,
- or a 'excel' temporary file or
- a known file type.

Firts let's define a simple fun that takes a list of fun and stop evaluating those fun as soon as a result is found:
% the simple case where the list is empty
any(_, []) -> undefined;

% the general case when the list contains funs...
any(Arg, [ {F, PrepareFun} | Funs] ) ->
        case F( PrepareFun(Arg) ) of
                undefined ->
                        any(Arg, Funs);

                _V ->
                        _V
        end.

In this code you'll notice that there are two fun()s:
- the 'F',
- the 'PrepareFun'.
The idea is that 'PrepareFun' will be called before calling 'F' to filter the argument 'Arg'.
Imagine that sometimes you need to extract the basename from the filename, or whatever else...

The code is a simple list iteration that recurse only if the result of the function call is 'undefined'.

Now that we have a valid fun that can iterate over a list of funs and stop whenever a valid result is found (or end of list), let's get back to our example, and build our 'filetype' function:
fileType(File) ->
        any( File, [ 
                        {fun word_temp/1, fun filename:basename/1}, 
                        {fun db_extension/1, fun lists:reverse/1},
                        {fun excel_temp/1, fun filename:basename/1}
                ]).

You can read this code like this:
"any of the funs from the list may determine the type of the file".
And once found, stops.

Let's describe those called functions 'db_extension/1', 'word_temp/1', 'excel_temp/1'...

First 'db_extension':
You'll notice that we test only the end of the filename, that's why the filename is
reversed before being passed to the function:
db_extension( "pmt."  ++ _ ) -> temp;
db_extension( "PMT."  ++ _ ) -> temp;
db_extension( "xcod." ++ _ ) -> doc;
db_extension( "cod."  ++ _ ) -> doc;
db_extension( "xslx." ++ _ ) -> xls;
db_extension( "slx."  ++ _ ) -> xls;
db_extension( "xtpp." ++ _ ) -> ppt;
db_extension( "tpp." ++ _ ) -> ppt;
db_extension( _ ) -> undefined.


The 'word_temp/1' need to call the basename of the file but we don't need the full path, so 'PrepareFun' is simply 'filename:basename/1' in this case:
word_temp( "~$"   ++ _) -> temp;
word_temp( "~WRD" ++ _) -> temp;
word_temp( "~WRL" ++ _) -> temp;
word_temp( _ ) -> undefined.


For 'excel_temp/1', the temp file is determined by a number written as 8 hexadecimal values. We use the re module to easily match this with the filename. In this case the 'PrepareFun' is also the 'filename:basename/1':
excel_temp( File ) ->
        ReList = [ <<"^[0-9A-Z]{8}$">> ],
        do_re(File, ReList).
        
% We are able to test many re but in the specific 
% case the list contains only one element...
do_re(_, []) -> undefined;
do_re(Subject, [ Re | Rest ]) ->
        case re:run(Subject, Re, [{capture,none}]) of
                nomatch ->
                        do_re(Subject, Rest);

                match ->
                        temp
        end.

From the re module, options "capture none" is used to only returns if the re match, and not the part that successfully match...
(this is simple optimisation, since we don't care about the matching part)

If we look at back at what we've done here, we can see that
fileType(File) ->
        any( File, [ 
                        {fun word_temp/1, fun filename:basename/1}, 
                        {fun db_extension/1, fun lists:reverse/1},
                        {fun excel_temp/1, fun filename:basename/1}
                ]).

can really easily extended with other functions, as long as those new functions take only one parameter...
fileType(File) ->
        any( File, [ 
                        {fun word_temp/1, fun filename:basename/1}, 
                        {fun db_extension/1, fun lists:reverse/1},
                        {fun excel_temp/1, fun filename:basename/1},
                        {fun firefox_temp/1, fun filename:basename/1},
                        {fun directory_temp/1, fun(X) -> X end}
                ]).

Conclusion:
Building list of functions is an efficient way of "testing many conditions".

Sunday, March 9, 2008

Directories Recursively and Simple Binary Matching

Hi, It has been a long time :)

So today two simple things first, some high order fun to work on tree:

-module(dir).
-export([create/1]).

create(List) ->
H = fun(X) ->
fun(Y) ->
Dir = filename:join([X,Y]),
io:format("mkdir(~s)~n", [Dir]),
Dir
end
end,
build(fun(X) -> X end, H, List).

build(_Fun, _Builder, []) ->
ok;
build(Fun, Builder, [Elem|List]) ->
NewFun = Builder(Fun(Elem)),
build(NewFun, Builder, List).


A sample session:

5> dir:create(["ab","cd", "ef","12", "35", "av"]).
mkdir(ab/cd)
mkdir(ab/cd/ef)
mkdir(ab/cd/ef/12)
mkdir(ab/cd/ef/12/35)
mkdir(ab/cd/ef/12/35/av)
ok

What's interesting is the Builder fun that construct the fun that will be called the next time.
That way we know were we are in the tree... The last Fun has all the knowledge of its ancestors :)

Now a simple binary matcher code, I need it while extracting values from regex results. It takes a list of offsets and returns what's inside:

slice(Slices, Bin) ->
slice(Slices, [], Bin).

slice([], Acc, _Bin) ->
lists:reverse(Acc);
slice([ {Start, Stop} | Rest ], Acc, Bin) ->
Len = Stop - Start,
<<_:Start/binary,Value:Len/binary,_/binary>> = Bin,
slice(Rest, [ Value | Acc ], Bin).


A sample session:

8> matcher:slice([{1,5},{8,10}], <<"this is a not a solution">>).
[<<"his ">>,<<"a ">>]


I hope that someone will find this valuable...

Tuesday, October 2, 2007

High Order Functions must be tested before use

From my previous article comments it is a small and efficient method of filtering datas:

Cpu = fun({cpu, _, _}) -> true; (_) false end.

But whenever we want to transposing it into a HOF (high order function):

Filter = fun(Elem) ->
fun({Elem, _, _}) -> true; (_) false end
end.

This naive approach doesn't work:

1> Filter = fun(Elem) -> fun({Elem, _, _}) -> true; (_) -> false end end.
#Fun<erl_eval.6.49591080>
2> C = Filter(cpu).
#Fun<erl_eval.6.49591080>
3> C({test, t, t}).
true % this should have been false ...

As explained in this document, we need to use guards to make our high order function effective:

The rules for importing variables into a fun has the consequence that certain pattern matching
operations have to be moved into guard expressions and cannot be written in the head of the fun.

The correct way is:

Filter = fun(Elem) ->
fun({X, _, _}) when X == Elem -> true; (_) false end
end.

So we must keep in mind that sometimes we should really check that our HOF is working as expected !

Monday, October 1, 2007

High Order Functions, filtering lists...

I have a list of collected cpu and network values, from eth0 and eth1:
 L =
[{cpu,user,<<"3.05">>},
{cpu,nice,<<"0.00">>},
{cpu,system,<<"0.72">>},
{cpu,iowait,<<"0.03">>},
{cpu,steal,<<"0.00">>},
{cpu,idle,<<"96.20">>},
{eth0,rxpck,<<"2.52">>},
{eth0,txpck,<<"0.15">>},
{eth0,rxbyt,<<"173.80">>},
{eth0,txbyt,<<"44.68">>},
{eth0,rxcmp,<<"0.00">>},
{eth0,txcmp,<<"0.00">>},
{eth0,rxmcst,<<"1.25">>},
{eth1,rxpck,<<"0.00">>},
{eth1,txpck,<<"0.02">>},
{eth1,rxbyt,<<"0.00">>},
{eth1,txbyt,<<"1.00">>},
{eth1,rxcmp,<<"0.00">>},
{eth1,txcmp,<<"0.00">>},
{eth1,rxmcst,<<"0.00">>}]

If I want to manipulate such data set I'll need some filter functions that will help me to extract values, I need cpu values and eth0 values. High order function can do that for me !

First we need to select tuples, 'cpu' tuples:

Cpu = fun({X, Y, Z}) ->
if X == cpu ->
true;
true ->
false
end
end.


A better and far more erlangish method (thanks Zvi):

Cpu = fun({cpu,_,_}) -> true;
(_) -> false
end.


Okay this fun will return true whenever the first element of the tuple is 'cpu'.
But this fun is static, since 'cpu' is written in the function body. Let's make it dynamic:

Filter = fun(Motif) ->
fun({X, Y, Z}) ->
if X == Motif ->
true;
true ->
false
end
end
end.


The new High order function 'Filter' is generated by 'fun(Motif)' and takes as argument a tuple '{X, Y, Z}', this is a fun that return a fun...
This function can be used like this:

List = {cpu, test, dummy}. % a sample list
Cpu = Filter(cpu). % generate the Cpu fun
Cpu(List). %executing the Cpu fun
true.
List2 = {test, cpu, dummy}. %Other dummy list
Cpu(List2).
false.


Come back to our initial data set, and realize that we must iterate thru the list to extract possible values. Iterate and apply a fun to every element is what we need to do, futhermore we need to retrieve matching values... In fact we need to build the list of extracted values, and this can be accomplished by 'lists:foldl':

extractor(Motif) ->
fun(L) ->
lists:foldl(
fun({X, Y, Z}, List) ->
if X == Motif
-> [ {Y,Z} | List ];
true -> List
end
end,
[], L)
end.


In details, 'List' is the accumulator list, the one that will grow with valid tuple, the one we will return. The fun is the same as described before. To make things clear 'extractor/1' returns a fun that will parse a list of tuple extracting values that matches 'Motif'.

Another easier method, using only list comprehension:

extractor(Motif) ->
fun(L) when is_list(L) ->
[ {Y, Z} || {X, Y, Z} <- L, X == Motif]
end.

Usage:

38> Cpu = module:extractor(cpu).
#Fun<module.1.131615259>
39> Cpu(L).
[{idle,<<"96.20">>},
{steal,<<"0.00">>},
{iowait,<<"0.03">>},
{system,<<"0.72">>},
{nice,<<"0.00">>},
{user,<<"3.05">>}]
40> Eth0 = module:extractor(eth0).
41> Eth0(L).
[{rxmcst,<<"1.25">>},
{txcmp,<<"0.00">>},
{rxcmp,<<"0.00">>},
{txbyt,<<"44.68">>},
{rxbyt,<<"173.80">>},
{txpck,<<"0.15">>},
{rxpck,<<"2.52">>}]


I use such code for my monitoring project, I use the sar output (sysstat package), and I graph using rrdtool... This is a part of a bigger project that will compose a 'scheduler'.

Saturday, September 29, 2007

Building Things with high order functions...

One thing that always surprise me, is code that extensively uses high order function or let's call that "functions that returns functions" (may be known also as "closure")...
What can I, myself, do with such things ?
After a little time reading and thinking, and reading and reading, I designed a really simple usage of all of this. I'll create some helper functions to build xml tags...

-module(tags).
-export([tags/1]).

tags(Elem) ->
fun(X) ->
"<" ++ Elem ++ ">" ++ X ++ "</" ++ Elem ++ ">"
end.


Let me explain:
  • Elem will be the enclosing tag
  • X is the parameter the function will receive at call time

The module in action:

1> c(tags).
{ok,tags}
2> Div = tags:tags("div").
#Fun<tags.0.28130594>
3> Div("html text").
"<div>html text</div>"
4>

'tags:tags' returns a function that will create div tags...

Now you're able to build a list of functions that will create valid output, without knowing the real syntax. You can see 'tags:tags' as a function that abstract the final notation of an element of your choice.

For example, building a 'title' tags is done by :

Title = tags:tags("title").

Building a list of functions for your language can be done like this:

lists:map(fun tags:tags/1, ["title", "div", "p", "ul", "li", "script"]).

Sticky